Atharwa Borkar

Data Engineer @IBM | Ex - Dentsu Global Services “Scalable Data Pipelines | Hadoop | HDFS Commands | Linux Commands | Apache Spark | Hive | Cloudera | SQL | AWS (S3, EMR, EC2) | Apache Airflow | Apache Iceberg”

Role
Data Engineer at IBM
Location
Mumbai, MH, IN
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Atharwa Borkar

I’m a Data Engineer with 4+ years of experience designing and delivering scalable, cloud-native data platforms that turn raw data into reliable, analytics-ready insights. I specialize in building high-performance batch and real-time data pipelines that support business intelligence, reporting, and decision-making at scale.My strength lies in working across AWS and Azure ecosystems, where I architect secure, cost-efficient solutions using services like Databricks, Redshift, Glue, Data Factory, Synapse, S3, and Blob Storage. I’ve built end-to-end ETL/ELT frameworks, optimized Spark workloads, and implemented data models that balance performance, governance, and usability.From a technical standpoint, I’m hands-on with Python, PySpark, SQL, and Spark, and experienced with Big Data technologies including Hadoop, Hive, and cloud data warehouses. I enjoy solving challenges around data reliability, scalability, and performance—especially in high-volume, fast-moving environments.I’m driven by ownership, clean architecture, and continuous learning—and I thrive in teams building modern, cloud-first data platforms.Open to Data Engineer roles. Let’s connect and build something impactful.

Experience

  1. Data Engineer

    IBM

    Nov 2025 — Present · Mumbai, IN

    Client - State Bank of India. Project - SBI Data LakeHouse Project.Project Role - PySpark Developer. Designed and supported a scalable data pipeline with Batch processing and Real time transactional data, ingesting from both Core Banking and Non-Core Banking Systems. The data storage architecture built on Apache Ozone, comprising of 3 layers-* Raw Source Layer - It is the layer in which the batch processing data from Core and Non-core Banking System is kept.* Enriched Layer - It is the layer in which the Real time transactional data is kept.* Conformed Layer - It is the layer in which the Final data is kept after applying all the transformations. Performed business-driven transformations using Higher level APIs(Dataframes, SparkSql) on Cloudera Data Platform (CDP), transforming raw data into analytically usable data. Stored transformed data in the Conformed Layer in Apache Iceberg format, enabling schema evolution, ACID transactions, and efficient analytical querying. Collaborated with the DataStage team, who managed and created Iceberg DDL tables in all the 3 layers for each pipeline maintaining proper Schema, allowing seamless direct writes of transformed data from PySpark into structured data in all the 3 Layers. Integrated real-time data ingestion using Kafka via GoldenGate, streaming data directly into the Enriched Layer. Developed and executed master-level user- defined PySpark transformations. Persisted finalized, analytics-ready datasets into the Conformed Data Layer, stored exclusively in Apache Iceberg format. Handled and processed petabytes of banking transactional data both Batch processing and Real time data with proper Spark optimization techniques making pipelines to run in more optimized way.

Education

  • University of Mumbai

    Bachelor's degree, Instrumentation engineering

    2017 — 2021

  • Aditya Birla Public School

    12th science, Physics/chemistry/maths

    2015 — 2017

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Atharwa Borkar — Data Engineer at IBM in Mumbai, MH, IN | Unifers