Mihir Shah
Principal Software Platform Engineer @Sony Pictures Networks India
Signup · Get unlimited contacts
WORK HISTORY
Principal Software Platform Engineer @Sony Pictures Networks India
Mumbai, IN
Led platform scalability initiatives enabling a live streaming platform to reliably support 10M+ concurrent users during the Asia Cup 2025, achieving 99.99% uptime under extreme peak traffic conditions. Drove a strategic modernization of high-traffic, mission-critical services by migrating legacy core systems to Amazon EKS, improving resilience and elastic scaling while delivering a 20% reduction in operational costs. Architected and owned a self-service, developer-first observability platform, automating alert onboarding from Coralogix to Datadog and reducing MTTx for critical incidents by 30%. Provided technical leadership and mentorship to 10 engineers, establishing best practices in cloud-native architecture, reliability engineering, and operational excellence across teams.
EDUCATION
Shri Baghubhai Mafatlal Polytechnic
Diploma, Computer Science Engineering
St. Lawrence High School
SSC
Dwarkadas J. Sanghvi College of Engineering
Bachelor of Engineering (BE), Computer Engineering
SKILLS
ABOUT MIHIR SHAH
I am a Principal Site Reliability Engineer dedicated to building and leading observability and reliability practices for some of the highest-traffic platforms in India. My career is defined by managing extreme scale - most recently handling 10 million concurrent users for live streaming at Sony Pictures Network India (SPNI) and previously 16 million concurrent users during the IPL at Dream11.I specialize in driving platform modernization, migrating legacy workloads to EKS, and architecting self-service observability tools that reduce MTTx metrics by 30%. I believe that reliability is a culture, and I am passionate about mentoring the next generation of engineers—having managed teams of 10+ junior engineers to bridge the gap between complex infrastructure and business stability.What I Bring to the Table:• Scale & Resilience: Proven track record of managing 5 GB/s logs throughput across 300+ components and ensuring high availability for tier-1 services.• Observability Strategy: Expert in automating alert onboarding and migrating massive monitoring environments (e.g, New Relic or Coralogix to Datadog).• Automation & IaC: Streamlining deployments using Terraform, Helm, and ArgoCD to create secure, HIPAA-compliant, for high-concurrency environments.• Incident Leadership: Leveraging AI tools for anomaly detection and automated RCA to cut post-incident analysis time by 35%.Continuous Learning: Driven by a commitment to technical evolution, I am currently exploring OpenTelemetry (OTEL), eBPF, and Cilium to push the boundaries of deep-system visibility and cloud-native networking.Core Tech Stack:• Cloud & Orchestration: AWS (EKS), GCP (GKE), Docker, Kubernetes.• Observability: Datadog, Coralogix, New Relic, OTEL, ELK Stack, Prometheus, Grafana.• DevOps/IaC: Terraform, Crossplane, ArgoCD, Helm, GitLab CI/CD.• Languages: Python, Go, Java, Bash.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.