Ganesh Makkina
Senior Data Engineer at S&P Global | ELT pipelines, Databricks, Kubernetes, Flask & FastAPI — building data infrastructure for financial analytics
- Role
- Senior Data Engineer at S&P Global
- Location
- New Brunswick, NJ, US
- LinkedIn followers
- 500 followers
About Ganesh Makkina
I build data systems that turn complex, high-volume financial workflows into reliable internal products teams can use independently.At S&P Global, I work on a pipeline that ingests, extracts, transforms, and standardizes thousands of SEC filings (10-K and 10-Q) for downstream analyst consumption. The system processes XBRL JSON filings, extracting millions of financial facts across hundreds of companies into Databricks Delta tables — running as scheduled production jobs with multi-threaded orchestration and automatic retry logic.The original pipeline was built on PySpark and Spark SQL in Databricks notebooks — heavy on cluster compute and tightly coupled to the Spark runtime. I led the effort to rewrite the extraction and transformation engine in pure Python, replacing Spark DataFrame operations with dictionary-based processing. The result: 90%+ reduction in per-file processing time, no Spark cluster dependency, and a pipeline that runs anywhere — Databricks, Docker, or Kubernetes — with zero code changes.I built a multi-stage validation framework that compares output across pipeline versions at every layer — master records, metadata, and financial facts — with normalization rules that separate real discrepancies from representation-only differences. The rewrite was validated row-by-row across tens of thousands of facts with exact parity confirmed.The transformation layer maps raw XBRL tags to standardized S&P mnemonics using a configurable rules engine with conditional logic, multi-column matching, and expression evaluation.On the productization side, I drove end-to-end delivery:1. Containerized with Docker (Flask API + FastAPI Sandbox UI)2. Deployed to Kubernetes on AWS EKS via Azure DevOps CI/CD with Helm charts3. Integrated HashiCorp Vault for zero-trust credential management4. Built a Sandbox UI for analysts to browse tables, compare mapping rules, trigger runs, and promote validated changes — no engineering support needed5. Designed for environment portability: all config via env vars, one Helm values file per environmentThe system now serves as the foundation for corporate XBRL financial data at S&P, with clean separation between extraction, transformation, and delivery.Core strengths: Data Engineering · Python · Databricks · SQL · ETL/ELT · Docker · Kubernetes · Flask · FastAPI · CI/CD · AWS EKS · Vault · System Design
Experience
Senior Data Engineer
Oct 2025 — Present · NY, US
Building a financial data pipeline that extracts, transforms, and standardizes SEC filings (10-K/10-Q) at scale, processing XBRL JSON filings and extracting millions of financial facts into structured Databricks Delta tables for downstream analyst consumption- Rewrote the extraction engine from PySpark/Spark SQL notebooks to a pure Python pipeline using dictionary-based processing — achieving 90%+ reduction in per-file processing time and eliminating the Spark cluster dependency entirely- Built a configurable transformation layer that maps raw XBRL tags to standardized S&P financial mnemonics using a rules engine with conditional logic, multi-column matching, and expression evaluation- Designed and implemented a multi-stage validation framework comparing pipeline output across versions at the master, metadata, and financial fact level — confirming row-by-row parity across tens of thousands of records- Deployed the full system to Kubernetes on AWS EKS via Azure DevOps CI/CD with Helm chart-driven configuration and HashiCorp Vault for secure credential injection- Built an internal Sandbox UI (FastAPI) that enables analysts to browse output tables, compare baseline vs. modified mapping rules, trigger pipeline runs on demand, and promote validated changes — removing the need for engineering involvement in day-to-day operations- Containerized the application with Docker (Flask health API + FastAPI UI), designed for environment portability with all configuration driven via environment variables and a single Helm values file per environment.
Education
Sri Gayatri Educational Institutions
10+2, MPC
Army Institute of Technology, Pune
Bachelor of Engineering, Computer Science
2017 — 2021
Rutgers University–New Brunswick
Master of Science - MS, Statistics - Data Science
Defence Laboratories School
10th, CBSE
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.