Anurag Ashu Mor

DevOps / MLOps Engineer @ Netomi | LLM Infra, Kubernetes, AWS, Ray Serve, vLLM | Cut inference cost 40% & saved $500K+/yr | ex-BluSmart, MX Player

Role
Devops Engineer Ii at Netomi
Location
Gurugram, HR, IN
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Anurag Ashu Mor

I make large language models cheaper, faster, and production-safe.In my last 13 months at Netomi — the agentic-AI platform powering customer experience for United Airlines, Delta, DraftKings and MetLife — I\'ve cut LLM inference cost 40%, saved $15K/month through MIG GPU partitioning, and reduced token latency 35% by rearchitecting serving on Ray Serve + vLLM. Before that, I led the EKS migration of 20+ microservices at BluSmart (Asia\'s largest EV-ride unicorn), halving deploy time and driving $500K+ in annual cloud savings.What I actually do day-to-day- Design GPU-dense Kubernetes platforms (EKS / GKE) for LLM, RAG and vector-DB workloads — Milvus, Qdrant, Pinecone- Serve & optimize Falcon-180B, Llama-3, BERT and GPT-family models via Triton, vLLM, Ray Serve and Kubeflow- Own the CI/CD and GitOps spine — Jenkins, ArgoCD, Terraform, Helm- Ship observability with Prometheus, Grafana, Datadog; on-call for Fortune-500 P0/P1s (OpenAI/Azure dependencies)- Write production Python + Go; published peer-reviewed research on AI-assisted surgical-skill assessmentI\'m exploring Senior / Staff DevOps, MLOps and AI-Platform roles at product companies solving hard infra problems at scale. DM me or email a••••••••@outlook.com. GitHub & paper in Featured below.#OpenToWork #MLOps #Kubernetes #AIInfrastructure

Experience

  1. Devops Engineer Ii

    Netomi

    Apr 2025 — Present · Gurugram, IN

    Netomi is the agentic-AI CX platform behind United, Delta, DraftKings, MetLife; 40K+ concurrent req/s, Y Combinator / Index / WndrCo-backed- Cut LLM inference cost 40%($X00K annualized) by migrating Falcon-180B and Llama-3 serving from vanilla HF Transformers to vLLM + Ray Serve on MIG-partitioned A100/H100 GPUs across EKS- Reduced p95 token latency 35%(1.8s → 1.17s) by tuning tensor-parallel sharding, KV-cache reuse and batching — directly improved SLA on airline-reservation agents handling 40K rps- Saved $15K/month in GPU spend via a MIG partitioning scheme co-locating inference + eval + embedding workloads on the same nodes- Led P0/P1 incident response for enterprise tenants (OpenAI/Azure upstream outages) — authored runbook that cut MTTR from 42→14 min- Own GitOps pipeline (ArgoCD + Jenkins + Terraform) across 60+ services spanning vector DBs (Milvus, Qdrant, Pinecone), LangChain orchestrators and Triton Inference Server endpoints.

Education

  • Government Senior Secondary School

    Senior-Secondary, Medical+Non-Medical +Computer Science

  • UPES

    Bachelor of Technology - BTech, Computer Science - DevOps

    2018 — 2022

  • The Sanskriti School

    Matriculation

    2014 — 2015

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Anurag Ashu Mor — Devops Engineer Ii at Netomi in Gurugram, HR, IN | Unifers