Alan Liu
Ph.D ML LLM
- Role
- Machine Learning Engineer at Amazon
- Location
- Seattle, WA, US
- LinkedIn followers
- 500 followers
About Alan Liu
I’m a Machine Learning Engineer focused on large-scale language models, distributed training systems, and reinforcement learning for alignment, currently working at Amazon in Seattle.My work centers on building and scaling foundation models across both pre-training and post-training (RL). I’ve led and contributed to end-to-end LLM systems, from model architecture and MoE optimization to large-scale GPU training and kernel-level performance tuning. This includes distributed training on clusters with thousands of GPUs, improving MFU through straggler mitigation and communication–computation overlap, integrating FP8 training, and delivering significant gains in training efficiency, memory usage, and inference speed.On the post-training side, I work on asynchronous reinforcement learning frameworks for LLM alignment, addressing real-world challenges such as training–inference mismatch, policy staleness, and large-scale rollout collection. I have hands-on experience with PPO-style algorithms, GRPO, and off-policy RL at production scale.Before Amazon, I was a Lead Software Engineer at Aptiv, where I designed and optimized high-performance, low-latency LSTM-based time-series models for automotive systems, achieving strong predictive accuracy while reducing inference latency from microseconds to nanoseconds through hardware acceleration and system-level optimizations.I hold a Ph.D. in Computer Engineering from the University of Michigan and have published 20+ peer-reviewed papers with 800+ citations, including first-author work at MLSys. My interests lie at the intersection of LLMs, systems, performance optimization, and scalable ML infrastructure.
Experience
Machine Learning Engineer
May 2024 — Present · Seattle, WA, US
My work centers on building and scaling foundation models across both pre-training and post-training (RL). I’ve led and contributed to end-to-end LLM systems, from model architecture and MoE optimization to large-scale GPU training and kernel-level performance tuning. This includes distributed training on clusters with thousands of GPUs, improving MFU through straggler mitigation and communication–computation overlap, integrating FP8 training, and delivering significant gains in training efficiency, memory usage, and inference speed.On the post-training side, I work on asynchronous reinforcement learning frameworks for LLM alignment, addressing real-world challenges such as training–inference mismatch, policy staleness, and large-scale rollout collection. I have hands-on experience with PPO-style algorithms, GRPO, and off-policy RL at production scale.
Education
Beihang University
Bachelor's degree, Electrical and Computer Engineering
2012 — 2016
University of Michigan
Ph.D, Computer engineer
2017 — 2022
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.