Vivek
Data Engineer · Lakehouse & AI-Ready Data Platforms · Governance · Databricks Certified · Azure | AWS | GCP. Spark
- Role
- Data Engineer at Samta.Ai
- Location
- Kollam, KL, IN
- LinkedIn followers
- 500 followers
About Vivek
I’m a Data Engineer with close to 4 years building data platforms that teams actually rely on ~ not just pipelines that run, but systems with governance, lineage, and quality baked in from day one. I’ve worked across banking, AML, and real estate - designing hybrid architectures that handle 600M+ records a month while staying compliant with GDPR, CCPA, and PCI DSS.My core work lives at the intersection of Databricks, on-premise Spark, and Delta Lake. I’ve built multi-tenant ingestion frameworks that cut new source onboarding from 3 days to under 2 hours, rolled out OpenMetadata as a unified data catalog across 270+ datasets, and built observability frameworks that helped teams go from firefighting incidents to resolving them in under 30 minutes.I don’t just move data. I make it AI-ready-clean, governed, and structured so ML models and AI agents can actually consume it without extra prep work.Outside work, I maintain an open-source Python package with downloads and write a technical blog read by data practitioners monthly.Right now I’m looking for senior data engineering roles focused on Databricks, data governance, or building AI-ready data platforms -on-premise, cloud, or hybrid.If that’s what your team is building, let’s talk. c••••••••@gmail.com
Experience
Data Engineer
Oct 2023 — Present · Noida, IN
Banking | AML | Transaction MonitoringBuilt a hybrid multi-tenant data platform handling 600M+ records/month (~2TB) across on-premise Spark clusters for banking tenants and Databricks on Azure for cloud tenants — all within GDPR, CCPA, and PCI DSS compliance boundaries.What I built:→ Reusable ELT ingestion framework across 270+ sources (PostgreSQL, MySQL, MSSQL, S3, GCS, Azure Blob). Cut new source onboarding from 3 days to under 2 hours→ 50+ table star schema data warehouse for AML investigations and transaction monitoring. Query times down 30% through partitioning, clustering, and join tuning→ Spark performance overhaul — fixed data skew, broadcast joins, partition pruning. Compute costs down 25%, job runtime down 30%→ OpenMetadata rollout as unified data catalog across all 270+ datasets — single source of truth for lineage, ownership, and schema→ Pipeline observability framework extended to AI agents for self-diagnosis and auto-remediation. MTTR dropped 40%, most incidents resolved under 30 minutes→ Real-time migration from overnight batch to Spark Structured Streaming — compliance teams got same-day data→ Full PII controls: masking, tokenization, encryption across all fields. Audited and signed off→ Automated ML lifecycle end-to-end: feature engineering → training → versioning → deployment via Docker and CI/CDReal Estate Data Platform · India | US | SingaporeSole owner of a daily pipeline processing 400K+ records across 3 geographies.→ Built and maintained 10+ production scrapers with Airflow DAGs for orchestration→ Raw CSV ingestion → Delta Tables on Databricks→ Zero unplanned downtime across the project lifetime despite frequent schema drift and website redesigns→ Data quality checks and reconciliation logic to catch duplicate listings before they reached reporting
Education
St Gregorios higher secondary school
secondary education, Biology science
2013 — 2015
kerala technological University
Bachelor of Technology - BTech, Mechatronics, Robotics, and Automation Engineering
2015 — 2019
National Institute of Electronics and Information Technology (NIELIT)
Advanced Diploma in Artificial Intelligence
2019 — 2020
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.