Ganesh Makkina

Senior Data Engineer at S&P Global | ELT pipelines, Databricks, Kubernetes, Flask & FastAPI — building data infrastructure for financial analytics

Role
Senior Data Engineer at S&P Global
Location
New Brunswick, NJ, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Ganesh Makkina

I build data systems that turn complex, high-volume financial workflows into reliable internal products teams can use independently.At S&P Global, I work on a pipeline that ingests, extracts, transforms, and standardizes thousands of SEC filings (10-K and 10-Q) for downstream analyst consumption. The system processes XBRL JSON filings, extracting millions of financial facts across hundreds of companies into Databricks Delta tables — running as scheduled production jobs with multi-threaded orchestration and automatic retry logic.The original pipeline was built on PySpark and Spark SQL in Databricks notebooks — heavy on cluster compute and tightly coupled to the Spark runtime. I led the effort to rewrite the extraction and transformation engine in pure Python, replacing Spark DataFrame operations with dictionary-based processing. The result: 90%+ reduction in per-file processing time, no Spark cluster dependency, and a pipeline that runs anywhere — Databricks, Docker, or Kubernetes — with zero code changes.I built a multi-stage validation framework that compares output across pipeline versions at every layer — master records, metadata, and financial facts — with normalization rules that separate real discrepancies from representation-only differences. The rewrite was validated row-by-row across tens of thousands of facts with exact parity confirmed.The transformation layer maps raw XBRL tags to standardized S&P mnemonics using a configurable rules engine with conditional logic, multi-column matching, and expression evaluation.On the productization side, I drove end-to-end delivery:1. Containerized with Docker (Flask API + FastAPI Sandbox UI)2. Deployed to Kubernetes on AWS EKS via Azure DevOps CI/CD with Helm charts3. Integrated HashiCorp Vault for zero-trust credential management4. Built a Sandbox UI for analysts to browse tables, compare mapping rules, trigger runs, and promote validated changes — no engineering support needed5. Designed for environment portability: all config via env vars, one Helm values file per environmentThe system now serves as the foundation for corporate XBRL financial data at S&P, with clean separation between extraction, transformation, and delivery.Core strengths: Data Engineering · Python · Databricks · SQL · ETL/ELT · Docker · Kubernetes · Flask · FastAPI · CI/CD · AWS EKS · Vault · System Design

Experience

  1. Senior Data Engineer

    S&P Global

    Oct 2025 — Present · NY, US

    Building a financial data pipeline that extracts, transforms, and standardizes SEC filings (10-K/10-Q) at scale, processing XBRL JSON filings and extracting millions of financial facts into structured Databricks Delta tables for downstream analyst consumption- Rewrote the extraction engine from PySpark/Spark SQL notebooks to a pure Python pipeline using dictionary-based processing — achieving 90%+ reduction in per-file processing time and eliminating the Spark cluster dependency entirely- Built a configurable transformation layer that maps raw XBRL tags to standardized S&P financial mnemonics using a rules engine with conditional logic, multi-column matching, and expression evaluation- Designed and implemented a multi-stage validation framework comparing pipeline output across versions at the master, metadata, and financial fact level — confirming row-by-row parity across tens of thousands of records- Deployed the full system to Kubernetes on AWS EKS via Azure DevOps CI/CD with Helm chart-driven configuration and HashiCorp Vault for secure credential injection- Built an internal Sandbox UI (FastAPI) that enables analysts to browse output tables, compare baseline vs. modified mapping rules, trigger pipeline runs on demand, and promote validated changes — removing the need for engineering involvement in day-to-day operations- Containerized the application with Docker (Flask health API + FastAPI UI), designed for environment portability with all configuration driven via environment variables and a single Helm values file per environment.

Education

  • Sri Gayatri Educational Institutions

    10+2, MPC

  • Army Institute of Technology, Pune

    Bachelor of Engineering, Computer Science

    2017 — 2021

  • Rutgers University–New Brunswick

    Master of Science - MS, Statistics - Data Science

  • Defence Laboratories School

    10th, CBSE

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Ganesh Makkina — Senior Data Engineer at S&P Global in New Brunswick, NJ, US | Unifers