Sudeep Govathoti
Data Engineer | Analytics Engineer | Python, PySpark, SQL, Azure Data Factory, Databricks, Synapse, AWS Redshift, Snowflake, dbt, Airflow | Building Observable, Production-Grade Data Pipelines
- Role
- Data Engineer at Verizon
- Location
- Charlotte, NC, US
- LinkedIn followers
- 500 followers
About Sudeep Govathoti
Data Engineer with 4+ years building production data platforms on Microsoft Azure (Synapse, Data Factory, Databricks, ADLS) and AWS (S3, Redshift) using Python, PySpark, SQL, dbt, and Apache Airflow. I specialize in dimensional modeling, incremental loads, partitioning strategy, schema evolution, data contracts, and CI/CD for pipelines — shipping observable data products backed by validation frameworks, pipeline monitoring, and clear data governance that unblock analysts and power executive-ready Power BI and Tableau reporting. What I\'ve delivered recently: • Cut pipeline failures by 40% and improved data accuracy by 30% through incremental loads and partitioning on Azure Synapse + Redshift. • Reduced pipeline runtime by 35% and cloud compute spend by ~20% via Spark tuning (partitioning, bucketing, broadcast joins, caching). • Cut time-to-detection of upstream data issues by 50% with anomaly detection and data observability patterns (alerts, freshness SLIs). Currently: Pursuing Azure Data Engineer Associate (DP-203), AWS Data Engineer Associate, and Databricks Data Engineer Associate certifications. Open to: Data Engineer / Analytics Engineer roles in Charlotte, NC (remote / hybrid / relocation). s••••••••@gmail.com
Experience
Data Engineer
Jul 2024 — Present
Cloud data pipelines on Microsoft Azure (Synapse Analytics, Data Factory, Databricks, ADLS Gen2) and Amazon Redshift for marketing, web analytics, and customer-behavior reporting.• Built end-to-end cloud data pipelines ingesting marketing, web-analytics, and customer-behavior data from 10+ enterprise sources into Azure Synapse Analytics and Amazon Redshift; incremental loads and partitioning strategy improved data accuracy by 30% and cut pipeline failures by 40%.• Developed ETL/ELT workflows in Azure Databricks using PySpark, Apache Spark, Python, and SQL, processing terabyte-scale datasets with aggregations, deduplication, and schema-evolution handling into curated Parquet data products for 4 downstream BI teams.• Automated orchestration with Azure Data Factory (ADF) and Apache Airflow DAGs; introduced CI/CD for data pipelines via Azure DevOps, eliminating 30% of manual intervention and tightening SLA adherence for operational and regulatory reporting.• Modeled star-schema data marts (fact + dimension tables) in Azure Synapse powering Power BI dashboards for marketing ROI, customer behavior, and operational KPIs; unlocked self-service analytics for 50+ users and saved ~15 analyst-hours per week.• Established data quality checks, data validation frameworks, data contracts, and pipeline monitoring with anomaly detection and data observability patterns (alerts, freshness SLIs), cutting time-to-detection of upstream issues by 50% and strengthening data governance and lineage.• Tuned Apache Spark jobs and SQL queries via partitioning, bucketing, broadcast joins, and caching, reducing pipeline runtime by 35% and cloud compute costs by ~20%.• Partnered with analysts, marketing, product, and platform engineers; rotated on-call to triage and resolve production incidents under strict SLAs.Tech: Python, PySpark, Apache Spark, SQL, Azure Synapse, Azure Data Factory, Azure Databricks, ADLS Gen2, Amazon Redshift, AWS S3, Power BI, Git, CI/CD, Azure DevOps, Agile/Scrum.
Education
Lindsey Wilson University
Master's degree
KL University
Bachelor's degree, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.