Mike Shen
Software Engineer @ AWS AI&ML | Builder, Writer and Speaker
- Role
- Software Engineer at Amazon Web Services (AWS)
- Location
- Bellevue, WA, US
- LinkedIn followers
- 500 followers
About Mike Shen
Infrastructure Engineer with 10+ years of experience architecting distributed systems for frontier-scale AI. Expert in optimizing the full ML stack—from HPC networking (NCCL/EFA) and custom CUDA kernels to multi-tenant Kubernetes schedulers. Proven track record of designing org-wide control planes for AWS SageMaker and building high-throughput inference engines from scratch.* ML Systems & HPC: FlashAttention-3, PagedAttention, vLLM Internals, Speculative Decoding, NCCL, EFA, RDMA, PyTorch FSDP.* Cluster Orchestration: Kubernetes (Custom Operators), Kueue, Ray, Slurm, Multi-Tenant GPU Scheduling.* Training & Alignment: SFT, DPO, RLVR, Distributed Checkpointing, Gradient Accumulation Strategies.* Core Infrastructure: Kafka, Spark, Redis, Lucene, Infrastructure-as-Code (Terraform/CDK).
Experience
Software Engineer
Jan 2021 — Present · Bellevue, WA, US
Next-Gen Compute Integration: Spearheaded the integration of revolutionary compute architectures (e.g, UltraServer GB200) into SageMaker, delivering up to 15x faster inference and 4x faster training performance. Personally architected foundational support for training jobs and advanced GPU rescheduling for high availability.LLM Alignment & Serverless Architecture: Architected high-scale serverless training and LLM alignment systems (SFT, DPO, RLVR, RLAIF) with dynamic Ray cluster spin-up, enabling model customization at scale.Distributed Systems & Latency Reduction: Led critical performance initiatives, including implementing Seekable OCI (SOCI) technology, reducing container cold-start latency by 75% and achieving a 6x improvement in image pull times in production (from 475s to 80s). Also drove platform consolidation to reduce P90 job latency from 5 seconds to 2 seconds.Kubernetes-Native Platform: Designed and implemented a custom Kubernetes-based scheduler (Kueue-based) for hierarchical resource fair-sharing across distributed GPU clusters (foundational to re:Invent launches like Crescendo) and consolidated AI pipelines, reducing end-to-end latency by 67%.
Education
Boston University
Master's degree, Computer Engineering
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.