Abhishek S.
Senior DevOps / SRE |Cloud AWS • GCP | Kubernetes & Terraform | CI/CD & GitOps | AIOps & Observability
- Role
- Senior Site Reliability Engineer at Altimetrik
- Location
- Gurugram, HR, IN
- LinkedIn followers
- 500 followers
About Abhishek S.
Senior DevOps and Site Reliability Engineer with 7+ years of experience running mission-critical production platforms across banking, telecom, and e-commerce environments. I specialize in designing resilient cloud-native systems on AWS, Azure, and GCP, with deep expertise in Kubernetes, Terraform, and GitOps-driven delivery.I have led observability and AIOps initiatives using Dynatrace, Prometheus, Grafana, and Splunk—driving measurable improvements in MTTR, alert quality, and incident prevention. My work spans progressive delivery strategies (blue-green and canary), release SLOs, automated rollback pipelines, and infrastructure standardization through IaC and configuration management.On the operations side, I focus on security-first platforms (IAM/RBAC, CIS hardening, vulnerability management), disaster recovery design, and multi-region resiliency. I am particularly interested in Internal Developer Platforms and agentic AI for operations—reducing toil while increasing engineering velocity.I thrive in environments where reliability, automation, and business outcomes intersect.
Experience
Senior Site Reliability Engineer
Jan 2025 — Present
Led production reliability and cloud operations for enterprise banking and industrial clients, driving observability modernization, release safety, and automation across Azure and Kubernetes platforms.Key ContributionsBuilt end-to-end observability for Azure workloads using Managed Prometheus, Grafana, and Azure Monitor, reducing MTTR by 25% and false alerts by 30%.Designed SLO/SLA dashboards and actionable alerting for AKS, PostgreSQL, and API services.Standardized Terraform and Ansible modules with approval workflows and guardrails to improve release safety.Acted as Build & Release–focused SRE owning CI/CD pipelines and production deployment strategies.Implemented blue-green and canary rollouts across EKS/GKE/AKS with HPA/VPA-driven capacity controls.Reduced rollback frequency and accelerated detection of faulty releases through post-deployment health checks.Automated patching and remediation using Ansible and event-driven workflows.Expanded Dynatrace and Splunk usage with business KPIs and noise reduction strategies.Introduced PagerDuty/Opsgenie runbooks and on-call playbooks to standardize incident response.Led cross-functional collaboration with product, security, and platform teams to improve operational readiness.
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.