Amit Agarwal

AI/ML Inference @Meta | Ex-Microsoft

Role
Software Engineer, Ml Inference at Meta
Location
Redmond, WA, US
LinkedIn followers
500 followers

About Amit Agarwal

Building cross-hardware ML Inference Stack @Meta for production ranking/recommendation workloads. Distributed inference of massive PyTorch ranking models on GPU, CPU and MTIA hardware targets, with a unified model processing, compiler + runtime stack and OpenAI Triton custom kernels. Accelerate time-to-production for ML innovations, with serving efficiency and reliability @Meta scale-I am a hands-on engineering leader with 15+ years of experience building, shipping, and operating AI/ML and HPC frameworks, platforms, runtimes, and services, at scale for enterprises and consumers. My passion is engineering distributed frameworks and platforms for compute intensive workloads, especially ML inference and training. I have built personalized Speech Recognition and Conversational AI services from ground up, designed and developed cross-hardware recommender/ranking ML execution backend stack and extensively worked on optimizing Speech and Computer Vision runtimes on GPUs and CPUs.• Leading development of vNext Ads ML PyTorch and OpenAI Triton based cross-hardware inference stack @Meta• Bootstrapped the ML cloud platform & infrastructure engineering team and systems @Modular (AI Infra Startup)• Built and led team that developed and operated personalized real-time Microsoft Teams meeting Speech Recognition services on Kubernetes, catering to millions of daily users and thousands of enterprises. Delivered industry leading operational cost-efficiency, service reliability and availability, and real-time latency SLAs • Founding member and dev lead of Microsoft’s deep-learning framework (CNTK) and a GPU platform for large scale distributed deep-learning training and inferencing• Architect and cross-org v-team lead at Microsoft for a massive scale batch (offline) processing system for cost-efficient Speech and Computer Vision AI inferencing workloads.• Led Microsoft Speech and OCR AI inference runtime optimizations utilizing novel quantization methods, GPU acceleration and CPU code generation for >10x price/performance improvements over 3 years

Experience

  1. Software Engineer, Ml Inference

    Meta

    Jun 2023 — Present · Bellevue, WA, US

    Building cross-hardware ML Inference Stack @Meta for production ranking/recommendation workloads. Distributed inference of massive PyTorch ranking models on GPU, CPU and MTIA hardware targets, with a unified model processing, compiler + runtime stack and OpenAI Triton custom kernels. Accelerate time-to-production for ML innovations, with serving efficiency and reliability @Meta scale.

Education

  • UCLA

    MS, Computer Science (Programming Languages and Systems)

  • National Institute of Technology Warangal

    BTech, Computer Science and Engineering

Skills

  • Algorithms
  • Computer Architecture
  • Multithreading
  • Parallel Computing
  • C++
  • Data Structures
  • C
  • High Performance Computing
  • Software Development
  • Perl
  • Java
  • Software Design
  • Scalability
  • Visual Studio
  • Unix
  • Debugging
  • Algorithm Design
  • Deep Learning
  • Machine Learning
  • Gpu Computing
  • High Performance Computing (Hpc)
  • Speech Recognition

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Amit Agarwal — Software Engineer, Ml Inference at Meta in Redmond, WA, US | Unifers