Amit Agarwal
AI/ML Inference @Meta | Ex-Microsoft
- Role
- Software Engineer, Ml Inference at Meta
- Location
- Redmond, WA, US
- LinkedIn followers
- 500 followers
About Amit Agarwal
Building cross-hardware ML Inference Stack @Meta for production ranking/recommendation workloads. Distributed inference of massive PyTorch ranking models on GPU, CPU and MTIA hardware targets, with a unified model processing, compiler + runtime stack and OpenAI Triton custom kernels. Accelerate time-to-production for ML innovations, with serving efficiency and reliability @Meta scale-I am a hands-on engineering leader with 15+ years of experience building, shipping, and operating AI/ML and HPC frameworks, platforms, runtimes, and services, at scale for enterprises and consumers. My passion is engineering distributed frameworks and platforms for compute intensive workloads, especially ML inference and training. I have built personalized Speech Recognition and Conversational AI services from ground up, designed and developed cross-hardware recommender/ranking ML execution backend stack and extensively worked on optimizing Speech and Computer Vision runtimes on GPUs and CPUs.• Leading development of vNext Ads ML PyTorch and OpenAI Triton based cross-hardware inference stack @Meta• Bootstrapped the ML cloud platform & infrastructure engineering team and systems @Modular (AI Infra Startup)• Built and led team that developed and operated personalized real-time Microsoft Teams meeting Speech Recognition services on Kubernetes, catering to millions of daily users and thousands of enterprises. Delivered industry leading operational cost-efficiency, service reliability and availability, and real-time latency SLAs • Founding member and dev lead of Microsoft’s deep-learning framework (CNTK) and a GPU platform for large scale distributed deep-learning training and inferencing• Architect and cross-org v-team lead at Microsoft for a massive scale batch (offline) processing system for cost-efficient Speech and Computer Vision AI inferencing workloads.• Led Microsoft Speech and OCR AI inference runtime optimizations utilizing novel quantization methods, GPU acceleration and CPU code generation for >10x price/performance improvements over 3 years
Experience
Software Engineer, Ml Inference
Jun 2023 — Present · Bellevue, WA, US
Building cross-hardware ML Inference Stack @Meta for production ranking/recommendation workloads. Distributed inference of massive PyTorch ranking models on GPU, CPU and MTIA hardware targets, with a unified model processing, compiler + runtime stack and OpenAI Triton custom kernels. Accelerate time-to-production for ML innovations, with serving efficiency and reliability @Meta scale.
Education
UCLA
MS, Computer Science (Programming Languages and Systems)
National Institute of Technology Warangal
BTech, Computer Science and Engineering
Skills
- Algorithms
- Computer Architecture
- Multithreading
- Parallel Computing
- C++
- Data Structures
- C
- High Performance Computing
- Software Development
- Perl
- Java
- Software Design
- Scalability
- Visual Studio
- Unix
- Debugging
- Algorithm Design
- Deep Learning
- Machine Learning
- Gpu Computing
- High Performance Computing (Hpc)
- Speech Recognition
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.