Hasif Subair

Senior Data Engineer @HelloFresh

Berlin, DE
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Feb 2022 — Present

Senior Data Engineer @HelloFresh

View department →

Berlin, DE

Stabilized Spark-based pipelines feeding global revenue dashboards, achieving a 97% reduction in failures and ensuring sub-hour data freshness for executive leadership.• : Orchestrated the migration of legacy data assets to modernized cloud systems with minimal downtime, focusing on long-term maintainability and reduced operational overhead.• & : Introduced robust quality gates and schema enforcement that increased asset reliability by 67%, significantly reducing downstream data drift and silent failures.• : Leveraged AWS EMR, PySpark, and Airflow to drive modular pipeline designs, improving test coverage and decoupling compute from storage for better cost-efficiency.• & : Won Gold in HelloFresh’s Data Olympics for leading a massive overhaul of the data catalog, eliminating redundant assets and streamlining discoverability across global engineering teams.

EDUCATION

2019 — 2019

Udacity

Data Engineering Nanodegree

2004 — 2008

University of Calicut

Bachelor of Technology - BTech, Electronics and Communications Engineering

SKILLS

Big DataStrutsJavaHbaseDatabasesElasticsearchMysqlMongodbSolution ArchitectureMicrosoft Sql ServerScalaSpring FrameworkApache KafkaCore JavaTestingJavascriptApache SparkWeb DevelopmentBusiness AnalysisJ2ee Application DevelopmentSqlLinuxHibernateJqueryAmazon Web Services (Aws)Agile MethodologiesWeb ServicesJava Enterprise EditionHtmlDroolsCssRequirements AnalysisSitemeshData AnalysisCassandraHibernate 3.1HadoopHiveSoftware DevelopmentNosql

ABOUT HASIF SUBAIR

I am a Senior Data Engineer with over 13 years of experience—the last 8 of which have been dedicated to building and scaling cloud-native data platforms on AWS and Azure. My career has spanned from the early days of the Hadoop ecosystem at TCS to architecting modern Lakehouse environments for companies like HelloFresh.I am currently focused on the infrastructure that makes AI actually useful at scale: governance through, high-performance storage with, and using frameworks like and to give AI agents the right data context. ’ : At Compredict, I led the migration from fragile, VM-based scripts to a Databricks stack on Azure. I built their first Medallion-style data lake, which didn\'t just make the data more reliable—it cut infrastructure costs by 50%: At HelloFresh, I rebuilt the Spark pipelines feeding global revenue dashboards. I introduced the data quality checks and schema enforcement needed to stop constant firefighting and ensure sub-hour data freshness for executive leadership: Lately, I’ve been architecting systems that automate the \"boring\" parts of the pipeline. I developed an engine that uses LLMs to perform automated schema discovery, generate data quality contracts, and deploy metadata-rich Delta tables directly to Unity Catalog via SQL Warehouse APIs: Solving data skew in Spark, architecting Lakehouse systems that don\'t break the bank, and figuring out how to build \"agent-ready\" data architectures using MCP: PySpark, Databricks (Unity Catalog, Delta Lake), AWS (EMR), Azure, Airflow, LangGraph, MCP, FastAPI, React.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Hasif Subair — Senior Data Engineer at HelloFresh in Berlin, DE | Unifers