Aditya Chauhan
Site Reliability Engineer | Kubernetes | Observability (Grafana, Splunk, ELK) | Production Reliability | MTTR Reduction | High Availability Systems
- Role
- Site Reliability Engineer at Tata Consultancy Services
- Location
- Gurugram, HR, IN
- LinkedIn followers
- 500 followers
About Aditya Chauhan
As a dedicated System Engineer / Production Support Executive with 5 years of experience in L1/L2 operations, I bring a strong blend of technical expertise and business understanding to ensure system stability, performance, and availability of mission-critical applications.With a Postgraduate Diploma in Business Management (PGDBM) from NMIMS and a bachelor’s in computer applications, I combine IT proficiency with managerial acumen—enabling me to lead teams, handle escalations, optimize workflows, and support strategic decision-making in fast-paced environments.I’ve successfully supported BFSI and Healthcare projects (Citibank, CVS Health), specializing in real-time monitoring, incident & problem management, RCA, and client communication using tools like ServiceNow, Splunk, AppDynamics, and ELK Stack. Recognized multiple times with Star of the Month/Quarter awards for KPI improvement, COB testing execution, and operational excellence.Open to remote or on-site opportunities in IT Production Support, IT Operations, or Team Lead roles, where I can leverage my skills, leadership experience, and commitment to delivering reliability and business continuity.
Experience
Site Reliability Engineer
Jan 2026 — Present · Gurugram, IN
Managing Kubernetes (k3s) production workloads including pods, services, deployments, and health validation• Handling controlled production deployments, rollout verification, and rollback coordination• Performing deep-dive alert investigation using Grafana dashboards and Splunk log correlation• Leading structured Root Cause Analysis (RCA) for recurring and high-impact incidents• Participating in P1/P2 incident bridges and reducing MTTR through systematic troubleshooting• Maintaining 99.9%+ uptime across critical enterprise applications• Executing Continuity of Business (COB) testing including controlled production failover, backup environment validation, and post-recovery stability monitoring• Managing enterprise batch orchestration using Control-M with dependency validation and recovery procedures• Coordinating with Development, Infrastructure, and Business teams during critical production events• Ensuring SLA/SLO compliance and proactive monitoring improvementsFocused on improving system reliability, strengthening observability practices, and enhancing production stability in 24x7 environments.
Education
NMIMS CDOE
Post Graduate Diploma Program, Marketing management
PSIT College of Higher Education
BCA - Bachelor of Computer Application, Computer and Information Sciences and Support Services
2017 — 2020
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.