Nan Cheng
Staff Data Engineer at H-E-B
- Role
- Staff Data Engineer at H-E-B
- Location
- Austin, TX, US
- LinkedIn followers
- 500 followers
About Nan Cheng
Certifications:GCP Skills Boost Camp Certificate: (07/2025)Databricks Certified Data Engineer Professional (07/2024)AWS Associate Developer Certified (10/2020)Personal Blog: http://nan0703.blogspot.com/Professional Summary:Data professional with over 10 years of hands-on experience in data analysis, data pipeline development, data modeling, database management, data visualization, and process automation. Proficient in Python and SQL, with expertise in data ETL processes and handling Big Data for analysis, extraction, normalization, and data governance; demonstrate a commitment to leveraging cloud technologies for scalable data solutions. Known for strong problem-solving abilities, attention to detail, and dedication to continuous learning and improvement. Excellent team player with effective communication skills and a deep passion for data-driven insights.Key Skills:Data AnalysisData Pipeline DevelopmentData ModelingDatabase ManagementData VisualizationProcess AutomationPythonSQLAWS Cloud TechnologiesCore Competencies:Strong analytical skills with a focus on deriving actionable insights from complex datasets.Proven ability to design and implement scalable data solutions for business needs.Dedication to staying updated with emerging technologies and industry trends.Effective communicator and team player, able to collaborate across multidisciplinary teams.
Experience
Staff Data Engineer
Oct 2022 — Present · Austin, TX, US
Effectively communicating with stakeholders and delivering Key Performance Indicators (KPIs) involves clarity, context, and alignment with organizational goals.* Data Processing Pipeline with Argo Workflows, Image registry(Harbor), Python, Databricks - Lead team to build scalable and resilient data processing pipelines using Argo Workflows on Kubernetes, leveraging its features for workflow orchestration, parallel processing, and monitoring.* Optimizing Spark Jobs for Cost Efficiency and Data Quality on AWS/Databricks - Lead project to integrate Spark optimization techniques with cloud cost reduction strategies and robust data quality/testing frameworks on AWS.* ETL Pipeline for Azure Data Integration, Event, Confluent Kafka - Lead project to establish a comprehensive CI/CD pipeline for Azure data integration and event processing leveraging Azure Data Factory, Event Hub, Event Grid, and Azure DevOps. By automating build, test, and deployment processes, organizations can achieve faster time-to-market, improved reliability, and scalability of data integration and event-driven applications in the Azure cloud environment.* Data Transformation and Analytics with dbt and Databricks - PoC project to integrate dbt with Databricks for building scalable and efficient data transformation pipelines in a cloud environment.* Machine Learning Model Deployment with AWS SageMaker, CodePipeline, and GitLab - PoC project to implement a robust CI/CD pipeline for deploying machine learning models using AWS SageMaker, AWS Lambda, and GitLab.
Education
Ohio University
M.S, Mathematics(Computational track, Computer Science option); Mathematics
2011 — 2013
Shandong Institute of Architecture and Engineering. Shandong
B.S, Information Management & Systems
2005 — 2009
Ohio University
M.S, Industrial & Systems Engineering
2009 — 2011
Skills
- Sas Programming
- Sas Certified Base Programmer
- Visio
- Data Analysis
- Pivot Tables
- Minitab
- Programming
- Statistics
- Microsoft Sql Server
- Kapow
- Matlab
- Databases
- Analysis
- Access
- Vba
- Access Database
- R
- Sas
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.