Shubham Mishra
Serving Notice Period | Sr. Data Engineer | Azure Databricks, PySpark, Snowflake, ADF | Building Scalable ETL/ELT Pipelines in Healthcare & Retail
- Role
- Senior Data Engineer at phData
- Location
- Indore, MP, IN
- LinkedIn followers
- 500 followers
About Shubham Mishra
I am a Senior Data Engineer with 4.5+ years of experience designing and building cloud data platforms on Azure for healthcare and retail organizations. I enjoy taking messy, siloed data and turning it into reliable, well‑modeled datasets that teams can actually use to make decisions.Most of my work sits at the intersection of Azure Data Factory, Azure Databricks (PySpark), Snowflake, and ADLS Gen2. I have built end‑to‑end ETL/ELT pipelines, from ingesting raw data into a Medallion (bronze–silver–gold) architecture to serving curated warehouse layers for analytics, reporting, and data science. Along the way, I focus on things that matter in production: data quality, observability, cost, and maintainability.On recent projects, I’ve implemented scalable PySpark workflows in Databricks, orchestrated with ADF/Airflow, to process high‑volume clinical and transactional data. This included standardizing schemas, enforcing business rules, and optimizing partitioning and joins so pipelines run predictably as data grows. In retail, I’ve supported customer and sales analytics by modeling star schemas in Snowflake and tuning SQL for faster dashboards and downstream models.I work best in environments where engineers talk directly to the business. I’m comfortable translating requirements from product owners, analysts, and domain experts into technical designs, and then iterating when needs change.Right now, I’m looking for Data Engineer roles where I can build or scale Azure/Snowflake‑based data platforms, work with modern tooling (PySpark, dbt, Airflow), and help teams move from “data is hard to trust” to “data is a dependable part of how we operate.”
Experience
Senior Data Engineer
Feb 2026 — Present · Bengaluru South, IN
Working on building and scaling modern cloud data platforms on Azure to support enterprise analytics, regulatory reporting, and near real-time insights.• Designed a scalable Medallion architecture (RAW → CLEAN → CONSUMED) on Azure using dbt and Snowflake, with the modelled layer serving as the enterprise data warehouse for healthcare analytics.• Migrated 500+ Oracle tables (TB-scale datasets) to Azure Data Lake (ADLS Gen2) and built dbt-based ELT pipelines to transform and load data into Snowflake. Currently leading the migration of 300+ additional tables to complete legacy system decommissioning.• Replaced legacy Informatica workflows with Airflow-orchestrated dbt pipelines, reducing end-to-end batch processing time from ~8–9 hours to approximately 2–3 hours while improving pipeline reliability.• Optimized Apache Airflow DAGs to support 200+ parallel jobs and scaled orchestration towards 800+ Azure Data Factory pipeline executions using CeleryExecutor.• Implemented centralized orchestration and monitoring using Airflow for Oracle → ADLS → Snowflake data flows, improving traceability, observability, and failure recovery across pipelines.• Collaborated closely with client stakeholders to gather requirements, provide delivery updates, and ensure data solutions aligned with healthcare compliance and reporting needs.Technologies: Python, PySpark, Snowflake, dbt, Azure Data Factory, ADLS Gen2, Apache Airflow, SQL
Education
Madhav Institute of Technology and Science, Gwalior
Bachelor's Degree, Mechanical Engineering
2018 — 2021
Rajiv Gandhi Proudyogiki Vishwavidyalaya
Diploma of Education, Mechanical Engineering
2015 — 2018
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.