Chandra Prakash Joshi
Principal Platform Engineer / Staff SRE | Cloud & Kubernetes Platforms | Reliability & Cost Optimization
- Role
- Senior Chief Engineer at Samsung Electronics
- Location
- Bengaluru, KA, IN
- LinkedIn followers
- 500 followers
About Chandra Prakash Joshi
Principal Platform Engineer / Staff SRE with 13+ years of experience designing, migrating, and operating highly reliable, cloud-native platforms across large-scale, latency-sensitive environments. I specialize in Kubernetes platform engineering, Site Reliability Engineering (SRE), and cloud infrastructure on AWS, with a strong focus on operability, cost efficiency, and platform sustainability.At Samsung, I lead platform modernization initiatives, including large-scale migration from Rancher-managed Kubernetes to Amazon EKS. I own platform reference architecture, rollout strategy, and operational standards across environments, enabling engineering teams through self-service platforms, golden paths, and standardized CI/CD pipelines using Terraform, ArgoCD, GitHub Actions, and Helm.I have extensive hands-on experience designing and operating highly available Kubernetes platforms across multiple AWS regions and hybrid environments. My work spans SRE practices such as defining SLIs/SLOs, error budgets, alert hygiene, incident management, and capacity planning, ensuring reliability without compromising latency.I architected centralized observability platforms using Prometheus, Grafana, Loki, and Tempo, implementing golden signal dashboards and distributed tracing with controlled telemetry overhead. By migrating from OpenSearch to Loki and tuning data ingestion pipelines, I delivered significant cost reductions while improving operational visibility.My background includes strong cloud cost governance and FinOps practices, resulting in multi-million-dollar annual savings through infrastructure right-sizing, storage optimization, automated environment shutdowns, and workload-aware scheduling strategies. I also enforce security best practices using RBAC, IAM/IRSA, Vault, and network policies to meet compliance and risk requirements.Earlier in my career, I worked with high-traffic consumer platforms such as Quikr, Goibibo, and Lenskart, where I managed large Linux-based infrastructures, built CI/CD pipelines, automated operations with scripting, and supported production systems handling significant scale and availability requirements.Beyond delivery, I actively mentor engineers, conduct knowledge-sharing sessions, participate in architecture discussions, and explore emerging tools and ideas through hackathons and proof-of-concepts. I am passionate about building reliable platforms that scale with business growth and reduce operational complexity.
Experience
Senior Chief Engineer
Mar 2020 — Present · Bengaluru, IN
As a Senior Chief Engineer/Staff SRE, I lead platform engineering and SRE initiatives for Samsung Ads, a highly latency-sensitive, real-time advertising platform operating at ~3 million RPS. I am responsible for designing, evolving, and operating Kubernetes and cloud platforms that support business-critical workloads where availability, scalability, and latency are core success metrics.I own the platform architecture and migration strategy for moving a hybrid infrastructure (AWS + self-managed data centers) running on Rancher-managed Kubernetes to a fully AWS-native Amazon EKS platform. This includes defining target architecture, phased rollout plans, automation standards, and risk-mitigation strategies to ensure stable migrations without sustained SLO impact.I design and operate highly available Kubernetes platforms across multiple AWS regions and availability zones, implementing workload isolation, topology-aware scheduling, and capacity planning to meet strict reliability and performance requirements. Infrastructure provisioning and lifecycle management are automated using Terraform, ensuring consistency and scalability.I drive SRE practices by defining SLIs, SLOs, and error budgets, improving alert quality, reducing operational toil, and strengthening incident response through standardized runbooks and post-incident reviews.I standardized CI/CD and deployment frameworks using Helm, GitHub Actions, and Jenkins, enabling safe, high-frequency releases via blue/green and canary deployments. I also architected centralized observability using Prometheus, Grafana, Loki, and Tempo, implementing golden-signal dashboards with controlled telemetry overhead.Through platform and observability optimization, infrastructure right-sizing, and automation, I delivered over $1.5M in annual cost savings. I enforce security best practices using RBAC, IAM, Vault, and network policies, and actively mentor engineers while contributing to cross-team architectural discussions.
Education
Dehradun Institute of Technology
Master of Computer Application (MCA), Computer Science
2009 — 2012
S.S.J Campus Almora
Bachelor of Computer Application (BCA), Computer Science
2006 — 2009
Skills
- Lvm
- Splunk
- Centos
- Nagios
- Linux
- Apache
- Jenkins
- Linux System Administration
- Haproxy
- System Administration
- Red Hat Linux
- Operating Systems
- Squid
- Solr
- Php
- Docker
- Mysql
- Network Administration
- Mongodb
- Network Security
- Zabbix
- Kubernetes
- Servers
- Troubleshooting
- Load Balancing
- Domain Name System (Dns)
- Git
- Python
- Ansible
- Varnish
- Shell Scripting
- Vpn
- Unix
- Ccna
- Software Deployment
- Network Load Balancing
- Networking
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.