Surabhi Shail
SDE @ Oracle Cloud | Distributed Systems · Cloud-native AI | Open to Senior SDE & AI Platform roles
- Role
- Software Engineer at Oracle
- Location
- Austin, TX, US
- LinkedIn followers
- 500 followers
About Surabhi Shail
Senior Software Engineer · 7 years building distributed systems and cloud-native infrastructure at Oracle Cloud.I specialize in fault-tolerant, highly available architectures at global scale — leading end-to-end delivery from design through production across 100+ cloud regions. At Oracle Cloud:Owned messaging platforms fanning out millions of pub/sub events monthly across email, SMS, webhooks, and cloud APIs. Designed retry architecture, migration strategies, and observability systems that keep delivery reliable at scale. What I work with:Distributed systems · Microservices · Event-driven architecture · Apache Kafka · Pub/SubJava · Python · Cloud-native development · Kubernetes · Docker · CI/CD · TerraformReliability engineering · Scalability · Observability · Grafana · IAMLangGraph · RAG pipelines · Agentic AI systems · LLM APIs · MCP🤖 Where I\'m heading:Cloud infrastructure taught me scale is only half the problem — the other half is making systems smart enough to handle what you didn\'t predict. Contributing to MCP-based DevOps automation, building agentic AI systems, using AI tooling daily. Systems that are not just scalable but intelligent. Open to Senior SDE and AI Platform Engineering roles. I write about distributed systems, production lessons, and system design tradeoffs from real engineering work.Latest: Kafka or a Queue? A Production Incident That Changed How I Think About System Designhttps://medium.com/@surabhishail.9/kafka-or-a-queue-a-production-incident-that-changed-how-i-think-about-system-design-5286a99e225eOthershttps://medium.com/@surabhishail.9/i-tried-building-real-agentic-ai-not-a-chatbot-heres-what-broke-first-2be47••••62b
Experience
Software Engineer
Nov 2020 — Present · Austin, TX, US
5+ years designing and running large-scale distributed systems in production — across dozens of regions, handling millions of events daily, where reliability isn\'t a feature, it\'s the job.* Rearchitected a high-throughput messaging pipeline to eliminate head-of-line blocking — a class of failure that caused latency spikes of 80x under load. Result: consistent delivery at scale, and room to grow 31% YoY without touching the architecture again.* Led a zero-downtime migration of a critical delivery service across all production regions — cut end-to-end latency from 60s to under 1s, modernised authentication, and eliminated an entire category of operational toil.* Designed and shipped end-to-end observability: dashboards, dimensioned metrics, automated availability reporting, and failure detection — so the team stops finding out about problems from customers.
Education
B.I.T Durg
Bachelor of Engineering (BE), Computer science and technology
2012 — 2016
Syracuse University College of Engineering and Computer Science
Master's degree, Computer Science
2018 — 2020
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.