Arpit Gaur

Staff SRE & Engineering Leader | GPU/ML Infrastructure at Scale | EKS • Kubernetes • Observability • Security | AWS • GCP • Azure

Role
Sde 4 at Adobe
Location
Pune Division, MH, IN
LinkedIn followers
500 followers

About Arpit Gaur

GPU training infrastructure breaks every assumption web services are built on. Jobs cannot be restarted. Nodes cannot be drained without cost. Patches cannot be applied without knowing what is in GPU memory. Reliability here requires a fundamentally different mental model — one I have spent my career building, and the last year applying across 3500+ GPU nodes at Adobe running uninterruptible AI training workloads. This constraint sharpens every decision. A wrong patching call kills training runs worth hours of GPU compute. A botched Kubernetes upgrade disrupts AI pipelines Adobe\'s products depend on. In the last year I rebuilt observability from the ground up, upgraded four EKS versions, designed a two-path security patching framework and established zero-trust networking — without disrupting a single training job. Technical depth alone does not deliver outcomes at this scale. Over 11 years across AWS, GCP and Azure, I have led teams of up to 12 through complete infrastructure transformations — Kubernetes adoption from zero, CI/CD for 100+ microservices, zero-downtime regional migrations at fintech scale. The best infrastructure leaders make architectural decisions with the credibility of someone who has debugged the system at 3am, and lead teams with the clarity of someone who understands exactly what they are asking engineers to build. The thread through my career — from being the first in my organisation to containerise a Windows application, to migrating node database clusters across five AWS regions with zero data loss, to operating one of the largest GPU Kubernetes fleets in the industry — is one consistent pattern: identifying where infrastructure needs to go before it gets there, and building it. The most dangerous infrastructure is the kind that appears to be working. Observability is not a dashboard — it is the difference between knowing your system is healthy and assuming it is. That conviction led AWS to approach me directly to co-author an official blog post on Amazon Managed Prometheus. I also contribute to the Grafana AMP plugin, used by engineers worldwide. Core stack: Kubernetes • EKS • AMP • Prometheus • Grafana • VPC Lattice • OPA • Kyverno • Karpenter • Terraform • Helm • Shell Scripting • Go • AWS • GCP • Azure Open to Staff, Principal SRE, Platform Engineering and Engineering Leadership roles — where technical depth, team leadership and infrastructure at scale need to coexist in the same person.

Experience

  1. Sde 4

    Adobe

    Apr 2024 — Present

    Responsible for reliability, observability and security of Adobe\'s AI/ML training infrastructure — one of the largest GPU Kubernetes deployments in the industry, comprising 2500+ reserved GPU nodes and on-demand GPU nodes across two EKS clusters in a single AWS account.• Resolved a critical observability gap — legacy Prometheus designed for 500–600 nodes was inadequately monitoring a 2500+ node GPU cluster. Architected full Prometheus upgrade integrated with Amazon Managed Prometheus (AMP), restoring complete observability across Adobe\'s AI/ML infrastructure.• Led EKS upgrade from 1.29 to 1.33 across 3500+ GPU nodes — planned 8–12 hour maintenance window and executed with zero unplanned downtime, delivering measurable AI/ML training performance improvements.• Designed a two-path security patching framework for 3500+ GPU nodes — live patching via a dedicated Kubernetes DaemonSet with dnf security updates, and AMI rebuild automation via Jenkins and Karpenter for kernel/driver-level vulnerabilities.• Implemented VPC Lattice across Kubernetes clusters — establishing zero-trust network architecture for Adobe\'s AI/ML workloads.• Built enterprise security and compliance posture from scratch — policy-as-code using OPA and Kyverno for Kubernetes admission control, ensuring CCF compliance and eliminating manual compliance overhead.

Education

  • SIIT, Pune

    Post Graduate Diploma, Advanced Computing

    2014 — 2015

  • College of Engineering and Technology, Bikaner

    Bachelor of Technology - BTech, Electrical, Electronics and Communications Engineering

    2008 — 2012

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Arpit Gaur — Sde 4 at Adobe in Pune Division, MH, IN | Unifers