Gandharv Pathak
Senior Data Engineer | Spark | Python | SQL | FinCrime
- Role
- Data Engineer 2 at NAB
- Location
- New Delhi, DL, IN
- LinkedIn followers
- 500 followers
About Gandharv Pathak
I am a Senior Data Engineer with 6+ years of experience building and leading large-scale data pipelines in financial services, currently working at National Australia Bank (NAB) in Gurugram as a Data Engineer 2 (Tech Lead, FinCrime & AML).At NAB, I serve as the cross-team Subject Matter Expert for Python, Spark, and Airflow across multiple FinCrime teams — authoring technical design documents, leading architecture reviews, and mentoring junior engineers. I designed and delivered a delta-based SCD2 batch ingestion system on Databricks that publishes 15M+ Kafka messages per day to downstream AML detection systems. One of my most impactful contributions was diagnosing and fixing a critical upstream defect that was causing Kafka volume to spike to 25M messages/day — a 3-day root cause investigation followed by a 2-day fix that resulted in a 96% reduction in Kafka I/O (from 9.12B to 372M messages/year) and cut job runtime from 4.2 hours to under 10 minutes. I also engineered PySpark pipelines processing 10M+ transactions/day across internet, mobile, and net banking channels, reducing end-to-end latency by 30% through partition optimisation and broadcast joins, and improved overall workflow operational efficiency by 40% through Airflow DAG design and SLA alerting.Prior to NAB, at RxLogix I built a Python REST API-based Generic Form Parser from scratch that achieved ≥80% extraction accuracy, eliminating manual data entry for enterprise pharma clients, alongside a high-availability Redis-based failover mechanism ensuring zero-downtime service continuity.At Iris Software, I replaced a legacy Excel-based financial system serving employees with a Python and PySpark platform for payroll, bonuses, and leave tracking, and built a Flask REST API backend for an internal HR chatbot also deployed to employees.Beyond my day job, I am an open-source contributor with a published PyPI package — dedupe-FuzzyWuzzy — a Python deduplication library, and a peer-reviewed research paper published in Springer with 100+ downloads. I hold a B.Tech in Computer Science from Jaypee Institute of Information Technology, Noida (2015–2019).
Experience
Data Engineer 2
Jan 2026 — Present · Gurugram, IN
At NAB, I serve as the cross-team Subject Matter Expert for Python, Spark, and Airflow across multiple FinCrime teams — authoring technical design documents, leading architecture reviews, and mentoring junior engineers. I designed and delivered a delta-based SCD2 batch ingestion system on Databricks that publishes 15M+ Kafka messages per day to downstream AML detection systems. One of my most impactful contributions was diagnosing and fixing a critical upstream defect that was causing Kafka volume to spike to 25M messages/day — a 3-day root cause investigation followed by a 2-day fix that resulted in a 96% reduction in Kafka I/O (from 9.12B to 372M messages/year) and cut job runtime from 4.2 hours to under 10 minutes. I also engineered PySpark pipelines processing 10M+ transactions/day across internet, mobile, and net banking channels, reducing end-to-end latency by 30% through partition optimisation and broadcast joins, and improved overall workflow operational efficiency by 40% through Airflow DAG design and SLA alerting.
Education
Jaypee Institute of Information Technology, Noida
Bachelor of Technology - BTech, Computer Science
2015 — 2019
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.