Sushil Yadav
Senior Data Engineer | Expertise in Apache Spark, Python, Scala, Azure (Data Factory, Databricks, Functions), AWS (S3, Redshift, Lambda), Snowflake, AIML | Passionate about Sustainability and Innovation
- Role
- Azure Data Engineer at Infosys
- Location
- Mumbai, MH, IN
- LinkedIn followers
- 500 followers
About Sushil Yadav
Senior Data Engineer with 3.5+ years of experience designing and building scalable, cloud-native data platforms on Microsoft Azure.I specialize in developing metadata-driven ETL frameworks and large-scale Spark processing systems using Azure Data Factory, Databricks, ADLS Gen2, Synapse Analytics, Python, SQL, and Delta Lake.At Infosys, I architected and owned a configuration-driven data platform based on the Medallion Architecture (Bronze–Silver–Gold), enabling low-code onboarding of new data sources and reducing change effort by ~50%. The platform currently processes 150+ data feeds and 100M+ daily records, with optimized Spark workloads that reduced runtime by 83% and improved cluster efficiency.My focus areas include:• Distributed data processing with Apache Spark• Performance optimization (data skew resolution, memory tuning, partition strategy)• Metadata-driven pipeline design• Data quality automation frameworks• Azure-based analytics architecturesI also lead and mentor a team of 4 engineers, conduct architecture reviews, and collaborate closely with product and business stakeholders to deliver production-ready, reliable data solutions across Healthcare and FinTech domains.Certified in:• Microsoft Azure Data Engineer (DP-203)• Microsoft Fabric Data Engineer (DP-700)• Databricks Data Engineer Associate
Experience
Azure Data Engineer
Jul 2024 — Present
Client: AmeriHealth Caritas (US Healthcare)• Designed and maintained large-scale healthcare data pipelines processing EDI 837 (Claims), 271 (Eligibility), and 820 (Payment/Remittance) transactions, supporting downstream analytics and reporting for payer operations.• Built ingestion and transformation frameworks for Facets core administration system data, integrating claims, membership, provider, and payment datasets into Azure-based Lakehouse architecture.• Developed robust ETL workflows to support Revenue Cycle Management (RCM) processes, ensuring accurate claim adjudication tracking, eligibility validation, and payment reconciliation.• Implemented data validation and reconciliation checks across 837 claims and 820 remittance files, improving financial data accuracy and reducing manual audit effort.• Processed high-volume healthcare transactional data (100M+ records) using PySpark and Delta Lake, optimizing Spark jobs to reduce runtime by 83% and improve cluster efficiency.• Designed schema mapping and transformation logic for HIPAA-compliant EDI structures, ensuring correct handling of segments (ISA, GS, ST) and claim-level hierarchies.• Collaborated with US business stakeholders and healthcare domain teams to translate payer reporting requirements into scalable data models and analytics-ready datasets.• Led a team of 4 engineers in supporting production healthcare data pipelines, managing deployments, incident resolution, and performance tuning.
Education
University of Mumbai
Bachelor of Engineering - BE
2018 — 2022
Guardian High School & Junior College
Secondary Education
2016 — 2018
Lokmanya Tilak College of Engineering
Bachelor's degree, Computer Engineering
2018 — 2022
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.