Vinay Kumar

Senior Site Reliability Engineer @redBus

Gurugram, HR, IN
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Oct 2014 — Present

Senior Site Reliability Engineer @redBus

View department →

Gurugram, IN

Improved platform reliability and reduced major incidents by ~35% through better alerting, RCA-driven improvements, and SLO-based monitoring.• Optimized AWS infrastructure (EC2, VPC, RDS, S3) to support rapid business growth, improving system performance and reducing cloud spend by 10–15%.• Designed and automated infrastructure using Terraform, cutting deployment time from hours to minutes and eliminating manual configuration drift.• Built and maintained CI/CD pipelines (Jenkins, Git), increasing release frequency and enabling safer, faster deployments.• Enhanced observability using Grafana, Zabbix, and CloudWatch, reducing MTTR by ~40% and improving proactive issue detection.• Managed on-call operations and led high-severity incident response, ensuring minimum downtime and driving long-term reliability improvements.• Supported containerized workloads using Docker and improved service scalability during peak traffic events.• Partnered with cross-functional teams (Dev, QA, Product) to drive SRE best practices, enhance system resilience, and deliver high-availability services to millions of users.

EDUCATION

2007 — 2013

Delhi University

Bachelor of Arts (B.A.), Human Resources Management and Services

2010 — 2014

Indian Institute of Technology, Delhi

Bachelor of Technology - BTech, Computer Engineering

2010 — 2012

Web Net Infotech

Diploma, Computer and Information Sciences and Support Services

2006 — 2007

SBV Chand Nagar

Intermediate, Economice, Mathematics, History, Pol Science, Hindi, English

2007 — 2008

Industrial Training Institute

Computer Programming, Specific Applications

2003 — 2004

Govt. Boys Sec School Chand Nagar

High School, Mathematics, Science, Social Study, Hindi, English

ABOUT VINAY KUMAR

Senior Site Reliability Engineer with 10+ years of experience in AWS cloud infrastructure, Linux administration, large-scale distributed systems, automation, observability, and production operations. Skilled in designing, deploying, and maintaining highly available, fault-tolerant, and scalable systems in fast-paced environments.Expert in AWS (EC2, VPC, RDS, S3, IAM, CloudWatch), Infrastructure as Code (Terraform), containerization (Docker), CI/CD pipelines (Jenkins, Git), and monitoring/alerting (Grafana, Zabbix). Strong background in performance optimization, log analytics, backup and recovery, capacity planning, and system hardening.Proficient in core SRE practices including SLIs/SLOs, error budgets, incident response, root cause analysis (RCA), on-call operations, change management, and service reliability improvement. Experienced in automating infrastructure, reducing operational toil, and ensuring consistent production stability.Key Skills: Site Reliability Engineering (SRE), AWS Cloud Architecture, Linux Systems Engineering, Terraform, CI/CD Automation, Docker, Monitoring & Observability, High Availability (HA), Disaster Recovery (DR), MySQL Administration, Cloud Security Best Practices, Networking (DNS, Load Balancing), Scalability Engineering, Troubleshooting & Debugging.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.