Rob Watson
VP / Sr Director Reliability & Cloud Ops | SLOs, Incident Lifecycle, Observability, FinOps/COGS | Regulated SaaS + Security
- Role
- Director, Site Reliability Engineering Production Support at Best Egg
- Location
- Middletown, DE, US
- LinkedIn followers
- 500 followers
About Rob Watson
I lead reliability and cloud operations for high-availability systems in regulated SaaS/fintech environments. My focus is practical: clear operating models, high-signal observability, disciplined incident execution, and cost/capacity governance that holds up as the platform scales.What I do:• Build and scale SRE / production operations organizations (including managers and multi-discipline teams)• Run incident lifecycle programs that produce follow-through (severity, roles, comms, postmortems, trend reviews)• Establish SLO/SLI/error budget practices to balance delivery velocity with stability• Own observability standards (Datadog + instrumentation expectations) and reduce alert noise• Drive FinOps/COGS governance (tagging/showback, guardrails, forecasting, optimization tradeoffs)• Partner with Security/Compliance to operationalize controls and maintain audit-ready operationsBackground: B.S. in Computer Networking & Security, early DoD infosec experience, and years in banking/fintech where security and compliance are part of day-to-day operations.I write about reliability leadership, observability that answers real questions, and practical automation that reduces toil and makes systems easier to run.
Experience
Director, Site Reliability Engineering Production Support
Mar 2020 — Present · Wilmington, DE, US
Director, Site Reliability Engineering / Production Support | Best EggLead a globally distributed SRE/Production Support team (5 direct reports) and drive org-wide reliability initiatives across Engineering, Infrastructure, and Security as part of an 8-person technology leadership council reporting to the CTOO. Built incident lifecycle standards, defined production readiness (“definition of done”), owned Datadog governance and cost controls, and delivered automation to reduce operational toil and improve response consistency.Optional bullets if you use bullets there:• Incident lifecycle (severity, roles, comms, postmortems, action tracking)• SLO/SLI practices + incident trend reviews + prevention tracking• Datadog standards + alert hygiene + instrumentation expectations• AWS/Datadog cost governance (tagging/showback, guardrails, optimization cadence)
Education
Wilmington University
Bachelor of Science, Computer Networking and Security
North Carolina Agricultural and Technical State University
Computer Science
Skills
- Switches
- Service Delivery
- Technical Support
- Business Process
- Vpn
- Vendor Management
- Information Technology
- Networking
- Major Incident Management
- Team Leadership
- Voip
- Data Center
- Microsoft Office
- Lan-Wan
- Management
- Software Installation
- Ip
- Business Process Improvement
- Computer Hardware
- IT Service Management
- Software Implementation
- Change Management
- Key Performance Indicators
- Troubleshooting
- Firewalls
- Routers
- Strategic Planning
- Project Management
- Help Desk Support
- Active Directory
- Hardware
- Customer Service
- Windows
- Fit/Gap Analysis
- Software Documentation
- Sdlc
- Security
- Problem Management
- Leadership
- Incident Management
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.