Sanjay B
Senior Data Engineer
- Role
- Senior Data Engineer at Healthfirst
- Location
- Buffalo, NY, US
- LinkedIn followers
- 500 followers
About Sanjay B
Experienced Data Engineer with 10+ years of delivering high-impact, data-driven solutions across diverse industries. I specialize in building scalable data pipelines, integrating data from heterogeneous sources, and enabling reliable analytics and machine learning workflows in cloud environments—especially on Google Cloud Platform (GCP). I have hands-on experience designing and automating ETL workflows using Cloud Data Fusion, BigQuery, and GCP-native services, which have significantly reduced development time and improved reusability across multiple projects. My work also includes extensive use of Spark, PySpark, and Spark SQL for large-scale data transformations and analytics. Skilled in Python, SQL, and R for data engineering, data wrangling, and performance optimization. I’ve developed robust pipelines using SSIS, automated workflows, and built end-to-end solutions for both structured and unstructured data. I\'m also experienced in implementing data quality and data governance standards, ensuring accuracy, compliance, and scalability. In addition to back-end engineering, I bring working knowledge of data visualization tools like Power BI, Tableau, and Looker, helping stakeholders make informed, data-driven decisions. I also develop RESTful APIs and data services using Python frameworks such as Flask and Django, with integration into both SQL (SQL Server, PostgreSQL, Oracle) and NoSQL (MongoDB) environments. I’m passionate about building resilient, production-ready data systems that drive business value, and I thrive in collaborative environments focused on continuous improvement, performance, and automation.
Experience
Senior Data Engineer
Jun 2024 — Present · NY, US
Working experience with distributed computing architectures like AWS cloud Platform(S3,EC2, Redshift, EMR, lambda, Glue, Elastic Search) Hadoop, Python, Spark.• Designed and developed ETL workflows using Informatica to extract, transform, and load data from heterogeneous sources into the enterprise data lake.• Created and scheduled SSIS packages for automated data integration from SQL Server and flat files into reporting and analytics platforms.• Utilized Power BI to create interactive dashboards and visualizations for executive reporting on patient data, improving decision-making timelines.• Collaborated with cross-functional teams to resolve complex ETL and data quality issues; documented resolution steps and communicated solutions clearly to stakeholders.• Conducted performance tuning on SQL Server queries and ETL jobs to optimize run-time and reduce load window by 35%.• Developed a data pipeline to ingest, process and deliver by batch processing and streaming data using spark, AWS EMR clusters, Lambda, and data bricks.• Implemented batch data processing, ETL, and ingestion into data warehouses using Lambda Python functions, Elastic Kubernetes Service (EKS), and S3 services through Airflow automation.• Ingested data into data lake (S3) and used AWS Glue to expose the data to redshift.• Provided ET solutions to migrate Teradata data from the on-premises system to AWS Red Shift and reduce run time by over 60%.• Used dbt (data build tool) to transform the data in Redshift after configuring the EMR cluster for data ingestion.
Education
Auburn University at Montgomery
Master's degree
Jawaharlal Nehru Technological University Hyderabad (JNTUH)
Bachelor's degree
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.