Sai Krishna Pala
Data Engineer | Azure Data Factory | Python & PySpark | SQL | ETL/ELT | Healthcare Data | DP-203
- Role
- Data Engineer at CVS Health
- Location
- Dayton, OH, US
- LinkedIn followers
- 500 followers
About Sai Krishna Pala
Data Engineer with hands-on experience designing and optimizing scalable ETL pipelines using Python, SQL, Apache Spark, and Microsoft Azure, with a strong focus on data quality, automation, and performance optimization.I specialize in building batch data pipelines, automating data ingestion using Azure Data Factory, and managing structured and semi-structured data in Azure Blob Storage and Azure SQL Database. My work supports analytics, reporting, and data science initiatives while ensuring compliance, reliability, and scalability.I bring practical experience across healthcare, manufacturing, and operations domains, along with strong collaboration in Agile/Scrum environments. I’m passionate about clean data, efficient pipelines, and continuously improving data systems that drive data-informed business decisions.Core Skills:Python (Pandas, PySpark)| SQL | Apache Spark | Azure Data Factory | Azure SQL | Blob Storage | ETL/ELT | Data Quality | Databricks | MongoDB | Agile
Experience
Data Engineer
Jul 2024 — Present · Pittsburgh, PA, US
At CVS Health, contributed to the design and optimization of scalable data platforms supporting large-scale healthcare analytics. Played a key role in building and maintaining robust ETL/ELT pipelines using Python, SQL, and AWS Glue, enabling seamless data integration from multiple enterprise systems into Snowflake and Amazon Redshift. Collaborated closely with data science teams to develop ML-ready datasets through feature engineering and data preprocessing, supporting predictive analytics initiatives. • Designed and automated scalable batch and near real-time data pipelines using Python and SQL for large healthcare datasets, improving processing efficiency by 25%. • Built and maintained ETL/ELT workflows using AWS Glue, S3, and Redshift for seamless data integration. • Developed ML-ready datasets by performing feature engineering and data preprocessing using Python (Pandas, NumPy). • Collaborated with data science teams to support machine learning model training and deployment workflows in Databricks and AWS environments. • Optimized pipeline performance and reliability, improving SLA adherence by 20%. • Implemented CI/CD practices using Git for version control and deployment automation. • Designed scalable data models for structured and semi-structured data to support analytics and ML use cases. • Implementeds data partitioning and performance tuning techniques in Redshift and Snowflake to optimize query execution. • Built reusable data pipeline frameworks to standardize ingestion and transformation processes. • Integrated REST APIs and third-party healthcare data sources into centralized data platforms. • Ensured data governance, security, and compliance standards (HIPAA-aligned practices) across pipelines. • Monitored production pipelines using CloudWatch and implemented alerting mechanisms.
Education
University of Dayton
Master's Degree, Computer Engineering
2023 — 2023
Aurora’s technological research institute
Bachelor of Technology
2018 — 2022
Avila University
Master's degree, computer science management
Aurora s Technological and Research Institute
Bachelor of Technology, Electrical, Electronics and Communications Engineering
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.