Ali Kazeroonian
Infrastructure & Fleet Platforms Software Executive | Capacity Delivery via Lifecycle Automation (Manufacturing→DC→Production) | AI/Accelerator Servers (NVIDIA, Trainium, Inferentia)
- Role
- Director, Software Development (Infrastructure Services) at Amazon Web Services (AWS)
- Location
- Seattle, WA, US
- LinkedIn followers
- 500 followers
About Ali Kazeroonian
I build and operate mission-critical B2B infrastructure services—from zero-to-one launches to global-scale operations—where hardware, software, manufacturing, and reliability intersect. Over the past 6+ years I’ve focused on delivering customer-ready hardware capacity by scaling software organizations, partnering with manufacturing and operations teams, and driving large, cross-functional programs.At AWS Infrastructure Services, I lead the software organization responsible for the global fleet platforms that test, provision, secure, and monitor AWS servers worldwide across the full lifecycle: manufacturing/assembly testing and firmware updates, data center provisioning, and in-production monitoring/diagnostics/security posture. My current scope is evenly split across provisioning, monitoring, manufacturing test/firmware, and security—with teams and mechanisms designed for high throughput and high reliability.Recent outcomes include reducing unsellable rates (servers failing testing) from >12 bps to <1 bp, while reducing accelerator server testing time by 40%. In 2025, my org also enabled rapid scale-out by adding 4 new manufacturing sites for assembling and testing accelerator systems.I’m a leader of leaders (orgs scaled from ~24 to 200+; 500+ hires across my career) with deep operating experience in services that must be reliable, secure, and fast-moving—especially in accelerated compute environments.
Experience
Director, Software Development (Infrastructure Services)
May 2019 — Present · Seattle, WA, US
I lead a multi-discipline organization delivering AWS’s global fleet platforms for GPU/accelerator and general compute server readiness across the full lifecycle: manufacturing/assembly (burn-in, test, firmware) → data center provisioning → in-production observability, security, and hardware diagnostics.Own platforms that drive fleet readiness and improve time-to-serve capacity, enabling service teams to detect hardware state, diagnose issues, and route remediation across the AWS fleet worldwide.Improved yield / sellable rate and throughput: reduced unsellables by servers in 2025 and cut accelerator server test time by 40% via automation and higher-fidelity signals.Lead a science/ML team building diagnostic models that recommend likely root cause and repair actions; reduced Hopper unsellables by 500+ in 4 months and reduced core compute server unsellables by using ML.Scaled manufacturing capacity: enabled rollout of 4 new accelerator assembly + test sites (2025) by standardizing tooling, test flows, and production playbooks.Drive execution across manufacturing partners, supply chain constraints, and data center operations to remove bottlenecks in the factory→DC→production pipeline and accelerate ready-to-serve hardware.Led teams responsible for fleet vetting, monitoring, and provisioning—ensuring hardware is validated, provisioned, and observable before entering customer-serving state. Built and operated services that improved fleet availability by accelerating detection/diagnosis and reducing escape rates through tighter validation signals. Partnered with infrastructure, hardware, and operations organizations to deliver multi-quarter programs. Built and scaled distributed teams; developed managers and senior engineers for larger ownership scope. Leader of leaders with budget ownership: scaled orgs from ~24 to 180+ and manage combined CapEx/OpEx $120M+/ year, while running global on-call/incident mechanisms and Correction of Error.
Education
Massachusetts Institute of Technology
Ph.D., Physics
1984 — 1991
Caltech
Bachelor of Science - BS, Physics
1982 — 1984
Skills
- Recruiting
- Mobile Applications
- Javascript
- Architecture
- Start-Ups
- Software Development
- Leadership
- Scalability
- Software Engineering
- Software Project Management
- Saas
- Representational State Transfer (Rest)
- Distributed Systems
- People Development
- Architectures
- Xml
- Rest
- Databases
- C#
- Product Management
- Product Development
- Web Applications
- Cloud Computing
- Hibernate
- Agile Methodologies
- Management
- Software as a Service (Saas)
- Java
- Entrepreneurship
- Microsoft Sql Server
- Tomcat
- Strategy
- Integration
- Web Development
- C++
- System Architecture
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.