Salman Ahmed
Senior Site Reliability Engineer (SRE) / Production Support Lead | P0/P1 Incident Management | Kubernetes | Splunk | Fortune 50 Environments
- Role
- Production Support Lead at Infosys
- Location
- Austin, TX, US
- LinkedIn followers
- 500 followers
About Salman Ahmed
Senior SRE / Production Support Lead with 10+ years in 24/7 production operations for large-scale systems (Fortune 50). I specialize in P0/P1 incident response, Kubernetes migrations, and observability with Splunk.I’ve led zero-downtime migrations of 30+ services, improved capacity from single digits to healthy headroom, and cut MTTR and repeat incidents by 30–40% through RCA and automation.Open to Senior SRE / Production Support roles where I can own reliability, incident management, and continuous improvement.
Experience
Production Support Lead
Oct 2022 — Present · Austin, TX, US
Leading production operations and site reliability engineering for mission-critical Apple services supporting millions of users. Responsible for incident management, Kubernetes migrations, CI/CD automation, and maintaining 99.9% SLA across 50+ microservices. & :Served as primary on-call DRI for P0/P1/P2 incidents across 50+ microservices, achieving 99.9% SLA and coordinating rapid triage with cross-functional teams :Reduced repeat incidents by 35% through systematic root cause analysis (RCA) and implementation of preventive controls using logs, metrics, and monitoring data :Led migration of 33+ critical services to Kubernetes-based platforms as Cluster Migration Lead, increasing capacity availability from 3% to 67% with zero downtime & :Built 5+ Splunk dashboards and automated alerts, reducing mean time to detection (MTTD) by 40% and enabling anomaly detection & :Implemented AI-powered automation for certificate management and operational workflows, reducing manual effort by 200 hours/month and improving response time by 45%/ & :Orchestrated CI/CD deployments for microservices and Apache Flink applications using Spinnaker, maintaining 99.5% deployment success rate with rollback readiness :Executed traffic shifting and failover procedures during maintenance windows and DR scenarios, minimizing customer impact to less than 5 minutes of downtime :Standardized operational runbooks and incident response procedures, reducing onboarding time for new engineers by 50%
Education
Alva's College of Education
Pre-University
2009 — 2011
NMAM Institute of Technology
Bachelor of Engineering - BE, Information science
2011 — 2015
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.