Shivam Pareek

Senior Data Engineer · Azure Databricks · PySpark · Delta Lake | Ex-TCS | Ex- Tiger Analytics

Role
Data Engineer at KPMG India
Location
Gurugram, HR, IN
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Shivam Pareek

I build data infrastructure that moves fast and breaks nothing.Over 5 years, I\'ve designed and shipped production-grade pipelines, lakehouses, and MLOps systems at scale — currently processing 3TB+ of daily financial data at KPMG India on Azure Databricks and Delta Lake.What I actually do:→ Design Medallion Architecture lakehouses (Bronze → Silver → Gold) from scratch, not templates→ Build PySpark pipelines optimized at the job level — partition tuning, broadcast joins, Z-ordering — not just \"working\" but fast→ Ship ML models to production with MLflow: cut deployment time 40%, reduced compute costs 30%→ Own the full stack: data ingestion → transformation → serving → monitoring → CI/CDRecent work I\'m proud of:• RAG-based Document Intelligence Pipeline using LangChain, Azure OpenAI, and pgvector — built end-to-end, handles enterprise-scale document retrieval• Real-Time Lakehouse on Azure Databricks with Unity Catalog governance across 20+ datasets, column-level security, full lineage• Retail forecasting platform at Tiger Analytics: 5M+ SKUs, 35% faster Spark jobs, 50% query improvement on Delta LakeStack: Python · PySpark · SQL · Azure Databricks · Delta Lake · ADLS Gen2 · ADF · MLflow · Unity Catalog · LangChain · pgvector · AWS (S3, Glue, Redshift, EMR) · Spark Streaming · CI/CD · GitI\'m selective about my next move — looking for senior data engineering or staff-level roles where the data problems are hard and the engineering bar is high.Open to: Google · Meta · Netflix · Amazon · or any team doing serious distributed systems work at scale. s••••••••@gmail.com

Experience

  1. Data Engineer

    KPMG India

    Jul 2025 — Present · Gurugram, IN

    Building production-grade data infrastructure at KPMG India — 3TB+ of daily financial data processed across a Medallion Architecture lakehouse on Azure Databricks.What I own:→ End-to-end MLOps platform: MLflow experiment tracking, model registry, automated deployment — cut time-to-production from weeks to days (40% faster)→ Unity Catalog governance: column-level security, data lineage tracking across 20+ datasets→ PySpark ETL/ELT pipelines for regulatory reporting and analytics — built for correctness, not just throughput→ Databricks cluster optimization: auto-scaling + spot instances → ~30% reduction in monthly compute spend→ CI/CD and DataOps: Git branching strategy + automated pipeline testing → 60% fewer deployment failures→ Cross-functional ownership: partnered with data science to productionise

Education

  • Sobhasaria Engineering College

    Bachelor of Technology - BTech, Mechanical Engineering

    2016 — 2020

  • PRINCE ACEDEMY OF HIGHER EDUCATION

    XII, Science

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Shivam Pareek — Data Engineer at KPMG India in Gurugram, HR, IN | Unifers