Shivam Pareek
Senior Data Engineer · Azure Databricks · PySpark · Delta Lake | Ex-TCS | Ex- Tiger Analytics
- Role
- Data Engineer at KPMG India
- Location
- Gurugram, HR, IN
- LinkedIn followers
- 500 followers
About Shivam Pareek
I build data infrastructure that moves fast and breaks nothing.Over 5 years, I\'ve designed and shipped production-grade pipelines, lakehouses, and MLOps systems at scale — currently processing 3TB+ of daily financial data at KPMG India on Azure Databricks and Delta Lake.What I actually do:→ Design Medallion Architecture lakehouses (Bronze → Silver → Gold) from scratch, not templates→ Build PySpark pipelines optimized at the job level — partition tuning, broadcast joins, Z-ordering — not just \"working\" but fast→ Ship ML models to production with MLflow: cut deployment time 40%, reduced compute costs 30%→ Own the full stack: data ingestion → transformation → serving → monitoring → CI/CDRecent work I\'m proud of:• RAG-based Document Intelligence Pipeline using LangChain, Azure OpenAI, and pgvector — built end-to-end, handles enterprise-scale document retrieval• Real-Time Lakehouse on Azure Databricks with Unity Catalog governance across 20+ datasets, column-level security, full lineage• Retail forecasting platform at Tiger Analytics: 5M+ SKUs, 35% faster Spark jobs, 50% query improvement on Delta LakeStack: Python · PySpark · SQL · Azure Databricks · Delta Lake · ADLS Gen2 · ADF · MLflow · Unity Catalog · LangChain · pgvector · AWS (S3, Glue, Redshift, EMR) · Spark Streaming · CI/CD · GitI\'m selective about my next move — looking for senior data engineering or staff-level roles where the data problems are hard and the engineering bar is high.Open to: Google · Meta · Netflix · Amazon · or any team doing serious distributed systems work at scale. s••••••••@gmail.com
Experience
Data Engineer
Jul 2025 — Present · Gurugram, IN
Building production-grade data infrastructure at KPMG India — 3TB+ of daily financial data processed across a Medallion Architecture lakehouse on Azure Databricks.What I own:→ End-to-end MLOps platform: MLflow experiment tracking, model registry, automated deployment — cut time-to-production from weeks to days (40% faster)→ Unity Catalog governance: column-level security, data lineage tracking across 20+ datasets→ PySpark ETL/ELT pipelines for regulatory reporting and analytics — built for correctness, not just throughput→ Databricks cluster optimization: auto-scaling + spot instances → ~30% reduction in monthly compute spend→ CI/CD and DataOps: Git branching strategy + automated pipeline testing → 60% fewer deployment failures→ Cross-functional ownership: partnered with data science to productionise
Education
Sobhasaria Engineering College
Bachelor of Technology - BTech, Mechanical Engineering
2016 — 2020
PRINCE ACEDEMY OF HIGHER EDUCATION
XII, Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.