Ali Kazeroonian

Infrastructure & Fleet Platforms Software Executive | Capacity Delivery via Lifecycle Automation (Manufacturing→DC→Production) | AI/Accelerator Servers (NVIDIA, Trainium, Inferentia)

Role
Director, Software Development (Infrastructure Services) at Amazon Web Services (AWS)
Location
Seattle, WA, US
LinkedIn followers
500 followers

About Ali Kazeroonian

I build and operate mission-critical B2B infrastructure services—from zero-to-one launches to global-scale operations—where hardware, software, manufacturing, and reliability intersect. Over the past 6+ years I’ve focused on delivering customer-ready hardware capacity by scaling software organizations, partnering with manufacturing and operations teams, and driving large, cross-functional programs.At AWS Infrastructure Services, I lead the software organization responsible for the global fleet platforms that test, provision, secure, and monitor AWS servers worldwide across the full lifecycle: manufacturing/assembly testing and firmware updates, data center provisioning, and in-production monitoring/diagnostics/security posture. My current scope is evenly split across provisioning, monitoring, manufacturing test/firmware, and security—with teams and mechanisms designed for high throughput and high reliability.Recent outcomes include reducing unsellable rates (servers failing testing) from >12 bps to <1 bp, while reducing accelerator server testing time by 40%. In 2025, my org also enabled rapid scale-out by adding 4 new manufacturing sites for assembling and testing accelerator systems.I’m a leader of leaders (orgs scaled from ~24 to 200+; 500+ hires across my career) with deep operating experience in services that must be reliable, secure, and fast-moving—especially in accelerated compute environments.

Experience

  1. Director, Software Development (Infrastructure Services)

    Amazon Web Services (AWS)

    May 2019 — Present · Seattle, WA, US

    I lead a multi-discipline organization delivering AWS’s global fleet platforms for GPU/accelerator and general compute server readiness across the full lifecycle: manufacturing/assembly (burn-in, test, firmware) → data center provisioning → in-production observability, security, and hardware diagnostics.Own platforms that drive fleet readiness and improve time-to-serve capacity, enabling service teams to detect hardware state, diagnose issues, and route remediation across the AWS fleet worldwide.Improved yield / sellable rate and throughput: reduced unsellables by servers in 2025 and cut accelerator server test time by 40% via automation and higher-fidelity signals.Lead a science/ML team building diagnostic models that recommend likely root cause and repair actions; reduced Hopper unsellables by 500+ in 4 months and reduced core compute server unsellables by using ML.Scaled manufacturing capacity: enabled rollout of 4 new accelerator assembly + test sites (2025) by standardizing tooling, test flows, and production playbooks.Drive execution across manufacturing partners, supply chain constraints, and data center operations to remove bottlenecks in the factory→DC→production pipeline and accelerate ready-to-serve hardware.Led teams responsible for fleet vetting, monitoring, and provisioning—ensuring hardware is validated, provisioned, and observable before entering customer-serving state. Built and operated services that improved fleet availability by accelerating detection/diagnosis and reducing escape rates through tighter validation signals. Partnered with infrastructure, hardware, and operations organizations to deliver multi-quarter programs. Built and scaled distributed teams; developed managers and senior engineers for larger ownership scope. Leader of leaders with budget ownership: scaled orgs from ~24 to 180+ and manage combined CapEx/OpEx $120M+/ year, while running global on-call/incident mechanisms and Correction of Error.

Education

  • Massachusetts Institute of Technology

    Ph.D., Physics

    1984 — 1991

  • Caltech

    Bachelor of Science - BS, Physics

    1982 — 1984

Skills

  • Recruiting
  • Mobile Applications
  • Javascript
  • Architecture
  • Start-Ups
  • Software Development
  • Leadership
  • Scalability
  • Software Engineering
  • Software Project Management
  • Saas
  • Representational State Transfer (Rest)
  • Distributed Systems
  • People Development
  • Architectures
  • Xml
  • Rest
  • Databases
  • C#
  • Product Management
  • Product Development
  • Web Applications
  • Cloud Computing
  • Hibernate
  • Agile Methodologies
  • Management
  • Software as a Service (Saas)
  • Java
  • Entrepreneurship
  • Microsoft Sql Server
  • Tomcat
  • Strategy
  • Integration
  • Web Development
  • C++
  • System Architecture

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Ali Kazeroonian — Director, Software Development (Infrastructure Services) at Amazon Web Services (AWS) in Seattle, WA, US | Unifers