Qinggang Zhou
Senior Staff Tl Tlm, Waymo Ml @Waymo
Signup · Get unlimited contacts
WORK HISTORY
Senior Staff Tl Tlm, Waymo Ml @Waymo
TL/TLM of the Runtime & Optimization team. Our team is responsible for the reliability and performance cross the model cycle, from model efficiency, quantization, sparsity, custom kernel libraries, MLIR/graph optimization, XLA, Triton, runtime, export, serving.* Spearheaded a transition in ML optimization strategy from late-stage compiler tuning to a proactive, holistic system-level approach. Transformed the Runtime & Optimization organization into a Center of Excellence.* Catalyzed \'0-to-1\' initiatives for Foundation Model acceleration, targeting both autoregressive and diffusion decoding. Achieved d a 1.4x speedup through speculative decoding and 2x through FFN sparsity.
EDUCATION
University of Washington
MSEE, EE
SKILLS
ABOUT QINGGANG ZHOU
ML Systems Leader with a proven track record of spearheading 0-to-1 AI initiatives and orchestrating complex, cross-company collaborations. Expert in the full-stack optimization of training and inference systems, spanning ML hardware (GPU/TPU), kernels, and compilers to ML frameworks. Deeply specialized in the system implications of ML workloads, including the latest models, foundation models, decoder efficiency, and diffusion models across both hyperscale data centers and edge environments.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.