Vinay Kumar
Senior Site Reliability Engineer @redBus
Signup · Get unlimited contacts
WORK HISTORY
Senior Site Reliability Engineer @redBus
Gurugram, IN
Improved platform reliability and reduced major incidents by ~35% through better alerting, RCA-driven improvements, and SLO-based monitoring.• Optimized AWS infrastructure (EC2, VPC, RDS, S3) to support rapid business growth, improving system performance and reducing cloud spend by 10–15%.• Designed and automated infrastructure using Terraform, cutting deployment time from hours to minutes and eliminating manual configuration drift.• Built and maintained CI/CD pipelines (Jenkins, Git), increasing release frequency and enabling safer, faster deployments.• Enhanced observability using Grafana, Zabbix, and CloudWatch, reducing MTTR by ~40% and improving proactive issue detection.• Managed on-call operations and led high-severity incident response, ensuring minimum downtime and driving long-term reliability improvements.• Supported containerized workloads using Docker and improved service scalability during peak traffic events.• Partnered with cross-functional teams (Dev, QA, Product) to drive SRE best practices, enhance system resilience, and deliver high-availability services to millions of users.
EDUCATION
Delhi University
Bachelor of Arts (B.A.), Human Resources Management and Services
Indian Institute of Technology, Delhi
Bachelor of Technology - BTech, Computer Engineering
Web Net Infotech
Diploma, Computer and Information Sciences and Support Services
SBV Chand Nagar
Intermediate, Economice, Mathematics, History, Pol Science, Hindi, English
Industrial Training Institute
Computer Programming, Specific Applications
Govt. Boys Sec School Chand Nagar
High School, Mathematics, Science, Social Study, Hindi, English
ABOUT VINAY KUMAR
Senior Site Reliability Engineer with 10+ years of experience in AWS cloud infrastructure, Linux administration, large-scale distributed systems, automation, observability, and production operations. Skilled in designing, deploying, and maintaining highly available, fault-tolerant, and scalable systems in fast-paced environments.Expert in AWS (EC2, VPC, RDS, S3, IAM, CloudWatch), Infrastructure as Code (Terraform), containerization (Docker), CI/CD pipelines (Jenkins, Git), and monitoring/alerting (Grafana, Zabbix). Strong background in performance optimization, log analytics, backup and recovery, capacity planning, and system hardening.Proficient in core SRE practices including SLIs/SLOs, error budgets, incident response, root cause analysis (RCA), on-call operations, change management, and service reliability improvement. Experienced in automating infrastructure, reducing operational toil, and ensuring consistent production stability.Key Skills: Site Reliability Engineering (SRE), AWS Cloud Architecture, Linux Systems Engineering, Terraform, CI/CD Automation, Docker, Monitoring & Observability, High Availability (HA), Disaster Recovery (DR), MySQL Administration, Cloud Security Best Practices, Networking (DNS, Load Balancing), Scalability Engineering, Troubleshooting & Debugging.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.