Gandharv Pathak

Senior Data Engineer | Spark | Python | SQL | FinCrime

Role
Data Engineer 2 at NAB
Location
New Delhi, DL, IN
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Gandharv Pathak

I am a Senior Data Engineer with 6+ years of experience building and leading large-scale data pipelines in financial services, currently working at National Australia Bank (NAB) in Gurugram as a Data Engineer 2 (Tech Lead, FinCrime & AML).At NAB, I serve as the cross-team Subject Matter Expert for Python, Spark, and Airflow across multiple FinCrime teams — authoring technical design documents, leading architecture reviews, and mentoring junior engineers. I designed and delivered a delta-based SCD2 batch ingestion system on Databricks that publishes 15M+ Kafka messages per day to downstream AML detection systems. One of my most impactful contributions was diagnosing and fixing a critical upstream defect that was causing Kafka volume to spike to 25M messages/day — a 3-day root cause investigation followed by a 2-day fix that resulted in a 96% reduction in Kafka I/O (from 9.12B to 372M messages/year) and cut job runtime from 4.2 hours to under 10 minutes. I also engineered PySpark pipelines processing 10M+ transactions/day across internet, mobile, and net banking channels, reducing end-to-end latency by 30% through partition optimisation and broadcast joins, and improved overall workflow operational efficiency by 40% through Airflow DAG design and SLA alerting.Prior to NAB, at RxLogix I built a Python REST API-based Generic Form Parser from scratch that achieved ≥80% extraction accuracy, eliminating manual data entry for enterprise pharma clients, alongside a high-availability Redis-based failover mechanism ensuring zero-downtime service continuity.At Iris Software, I replaced a legacy Excel-based financial system serving employees with a Python and PySpark platform for payroll, bonuses, and leave tracking, and built a Flask REST API backend for an internal HR chatbot also deployed to employees.Beyond my day job, I am an open-source contributor with a published PyPI package — dedupe-FuzzyWuzzy — a Python deduplication library, and a peer-reviewed research paper published in Springer with 100+ downloads. I hold a B.Tech in Computer Science from Jaypee Institute of Information Technology, Noida (2015–2019).

Experience

  1. Data Engineer 2

    NAB

    Jan 2026 — Present · Gurugram, IN

    At NAB, I serve as the cross-team Subject Matter Expert for Python, Spark, and Airflow across multiple FinCrime teams — authoring technical design documents, leading architecture reviews, and mentoring junior engineers. I designed and delivered a delta-based SCD2 batch ingestion system on Databricks that publishes 15M+ Kafka messages per day to downstream AML detection systems. One of my most impactful contributions was diagnosing and fixing a critical upstream defect that was causing Kafka volume to spike to 25M messages/day — a 3-day root cause investigation followed by a 2-day fix that resulted in a 96% reduction in Kafka I/O (from 9.12B to 372M messages/year) and cut job runtime from 4.2 hours to under 10 minutes. I also engineered PySpark pipelines processing 10M+ transactions/day across internet, mobile, and net banking channels, reducing end-to-end latency by 30% through partition optimisation and broadcast joins, and improved overall workflow operational efficiency by 40% through Airflow DAG design and SLA alerting.

Education

  • Jaypee Institute of Information Technology, Noida

    Bachelor of Technology - BTech, Computer Science

    2015 — 2019

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Gandharv Pathak — Data Engineer 2 at NAB in New Delhi, DL, IN | Unifers