Mayur Jain
Data Engineer | Python • SQL • Airflow • Snowflake • Databricks | LLM/Agentic Pipeline Optimization | MS CSULB | 5+ YOE
- Role
- Data Engineer at Saayam For All
- Location
- Long Branch, NJ, US
- LinkedIn followers
- 500 followers
About Mayur Jain
I’m a Data Engineer with 5+ years of experience designing, building, and optimizing large-scale data ecosystems across healthcare, consulting, and finance. My work focuses on creating production-grade data pipelines, scalable architectures, and ML/AI-driven solutions that deliver measurable business impact—not just theoretical models.I’ve engineered end-to-end data lakes, automated ingestion frameworks, quality monitoring systems, and predictive analytics pipelines using technologies like AWS (S3, Glue, Redshift, Lambda, EMR), Spark, Python, SQL, Kafka, Airflow, and Snowflake. My experience spans high-volume healthcare claims, financial datasets, and real-world evidence (RWE/RWD), enabling analytics teams to make faster, data-backed decisions.I’m passionate about bridging the gap between engineering and data science. I’ve deployed machine learning models into production using SageMaker, built RAG/LLM pipelines for unstructured text search, and implemented automated monitoring to ensure reliability, compliance, and performance. Whether it\'s optimizing a Spark job, improving data reliability, or building dashboards that uncover actionable insights, I focus on delivering solutions that are scalable, maintainable, and aligned with business goals.I enjoy working in collaborative, fast-paced environments where data drives strategy. My goal is to continue building powerful data platforms and AI-driven applications that solve complex problems and improve real-world outcomes.Always open to connecting with other data professionals, innovators, and organizations leveraging data to create value.
Experience
Data Engineer
Jul 2025 — Present · US
Built a unified healthcare Data Lake on AWS (S3, Glue, Lambda, Redshift, EMR) to consolidate fragmented claims, EHR, and policy data, improved data availability for analytics teams by 60%, and eliminated multiple legacy ingestion jobs.•Developed predictive models in Python/R to detect high-risk members, early chronic-condition indicators, and cost anomalies, integrating these models into operational workflows via SageMaker endpoints with automated model monitoring.•Implemented RAG pipelines using LLMs and vector search to extract insights from policy documents, clinical guidelines, and claims notes, reducing manual review time for care management teams by 55%.•Applied machine learning techniques on large-scale healthcare datasets to support risk stratification and outcome prediction, collaborating closely with data science teams to operationalize models in production environments.•Built and automated data quality frameworks for PHI/PII compliance (HIPAA) using Python, PyDeequ, and CloudWatch, catching schema drift, duplicate claims submissions, and data corruption before reaching downstream systems.•Created dashboards in Power BI/Tableau for Claims Ops and Care Management, visualizing utilization trends, fraud patterns, reimbursement cycles, and provider performance; enabled leadership to identify >$8M in avoidable spend annually.
Education
Guru Tegh Bahadur Institute Of Technology
Bachelor of Technology (B.Tech.), Information Technology
California State University, Long Beach
Master of Science - MS, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.