Harpal Singh
Lead Data Engineer @Capgemini
Signup · Get unlimited contacts
WORK HISTORY
Lead Data Engineer @Capgemini
CA
Requirement gathering, technical feasibility, brainstorming and user stories creation.• Develop data pipelines from scratch based on the client requirements using AWS, Python, APIs, Bigdata Hadoop, Hive, Pyspark, Databricks, SQL, Snowflake, Postgres, Airflow.• Developed application to sync domain hierarchy and domain table association data.• Developed metadata copy utility to copy metadata from source to target table/views.• Created stewardship utility to sync stewards’ data between domain, table and column levels.• Created data source onboarding utility to automate the manual process.• Developed a POC to migrate snowflake data pipeline to Pyspark.• Developed new pyspark codes to replace existing Snowflake programs and do Data Validation.• Lead the team and coordinate with client to discuss technical challenges and provide solutions. Delta logic: Implemented logic to copy incremental data instead of full data refresh which reduced the job execution time by 2/3rd and saved infrastructure cost. Significant reduction in compute cost after migrating data pipeline from snowflake to pyspark.
ABOUT HARPAL SINGH
Databricks Certified Lead Data Engineer with 12+ years of experience in designing and…
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.