Vinay Kumar
Site Reliability Engineer (SRE) | AWS | OCI | Terraform | Docker | Kubernetes | Incident & On-Call Management | Monitoring & Automation | Cloud Infrastructure
- Role
- Service Reliability Engineer at Oracle
- Location
- Bengaluru, KA, IN
- LinkedIn followers
- 500 followers
About Vinay Kumar
Site Reliability Engineer (SRE) with 4+ years of experience in cloud operations, automation, and large-scale production environments across AWS and Oracle Cloud. I specialize in improving reliability, reducing MTTR, automating manual operations, and building resilient systems using DevOps and SRE principles.I have led 50+ high-severity incidents, delivered 99.98%+ service availability, and automated 60% of operational tasks, significantly reducing toil and improving on-call efficiency.My expertise spans AWS, OCI, Terraform, CI/CD, Docker, Kubernetes, Prometheus, Grafana, and infrastructure automation.Core Strengths:• Incident Management, RCA, War Room Leadership• AWS Infrastructure & Terraform Automation• Monitoring, Logging & Observability (Prometheus / Grafana / CloudWatch)• CI/CD Pipelines (Jenkins, GitHub)• Linux Administration, Shell Scripting• Cloud Security (IAM, SSL/TLS, Network Security)• Scaling, Load Balancing, SLO/SLI/SLA-based operationsI’m passionate about delivering reliable systems, reducing operational friction, and collaborating with engineering teams to drive long-term stability and efficiency.If you’re building DevOps/SRE teams or want to collaborate, feel free to connect.
Experience
Service Reliability Engineer
May 2024 — Present · Bengaluru, IN
Led 50+ high-severity incident responses across global SaaS regions (JAPAC, EMEA, AMER), ensuring minimal customer impact and seamless stakeholder coordination-Improved platform uptime to 99.98%+ by implementing automated incident detection, RCA-driven fixes, and proactive reliability measures-Enhanced monitoring coverage by 35%, reducing MTTD & MTTR by 40% through optimized alerting, dashboards, and observability improvements-Conducted 100+ RCA & post-incident reviews, implementing preventive actions that reduced repeat incidents by 35%-Automated 60% of operational runbooks using Shell scripting and tooling, saving 20+ engineer hours per month and reducing on-call toil-Strengthened compliance and security posture aligned with SOC 2, ISO 27001, and internal Oracle cloud governance standards-Delivered real-time MI dashboards, outage reports, and strategic insights directly to VP-level leadership, improving decision-making during critical events-Collaborated with engineering, PM, and ops teams to ensure service continuity during major upgrades, deployments, and migrations- Standardized SOPs and runbooks, improving cross-regional operational consistency and onboarding efficiency.
Education
BIT Sindri
Bachelor's of tecnology
2017 — 2021
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.