Praveen Thunga
Site Reliability Engineer | Observability & Kubernetes Platform | Dynatrace | MuleSoft Runtime Fabric | Azure
- Role
- Site Reliability Engineer at Edward Jones
- Location
- Dallas-Fort Worth, TX, US
- LinkedIn followers
- 500 followers
About Praveen Thunga
Site Reliability Engineer | 13+ Years SRE/Observability | Dynatrace | Azure AKS | MuleSoft RTFDriving observability-first SRE at Edward Jones: reduced alert noise ~35-40%, incident triage ~25-30%, standardized K8s/service dashboards for SRE/NOC/leadership.Core Expertise:• Dynatrace (SaaS/Managed): DQL, synthetic monitoring, problem analysis, alerting profiles• Azure AKS: cluster health, pod restarts, node pressure, capacity planning • MuleSoft RTF: API latency, policy validation, integration reliability• Alert engineering: anomaly detection, false positive reduction (~30%)• Incident automation: Dynatrace → ServiceNow/Jira workflows13-Year Progression:Linux Admin → DevOps → SRE (Edward Jones, Equifax, AmEx, FedEx)Certifications: AZ-104 | Dynatrate Certified | AWS Solutions ArchitectOpen to remote US SRE/Observability roles.H4 EAD- eligible for any US employer, no sponsorship needed.
Experience
Site Reliability Engineer
Feb 2023 — Present · US
Led observability architecture and implementation across business-critical platforms on Azure AKS, using Dynatrace SaaS to deliver consistent system health visibility.• Designed and standardized dashboards covering services, APIs, Kubernetes workloads, and infrastructure, now used during incidents, operational reviews, and leadership reporting.• Engineered alert quality improvements by tuning thresholds, anomaly detection, and signal routing, reducing false positives by ~30% and improving on-call response.• Accelerated root-cause analysis by leveraging Dynatrace problem cards, DQL queries, and distributed traces across microservices and integrations.• Monitored AKS cluster health including pod restarts, node pressure, memory over-commit, and CPU saturation trends, identifying risks before production impact.• Strengthened MuleSoft Runtime Fabric (RTF) reliability by monitoring API latency, policy validation failures, and integration health across gateway and application layers.• Implemented synthetic monitoring (browser and HTTP clickpaths) to proactively detect availability and performance issues from an end-user perspective.• Built custom metrics, key requests, conversion goals, and anomaly detection rules, improving early detection of degradation and SLO tracking.• Defined tagging strategies, management zones, and alerting profiles, aligning observability with team ownership and platform boundaries.• Integrated Dynatrace with ServiceNow and Jira, automating incident creation and improving response workflows.• Led incident response and post-incident analysis, delivering RCA findings and telemetry-backed recommendations to prevent recurrence.• Partnered closely with MuleSoft platform teams, application owners, and leadership, translating telemetry into actionable reliability improvements.
Education
Sri Venkateswara University
Bachelor's degree, Computer Science
Adarsh Management Institute of India
Master of Business Administration - MBA
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.