Mohammad Khan
“Experienced Data Architect | Data Engineer | Big Data | Cloud (Azure,AWS) | Spark, Scala, Python | Data Pipelines | Machine Learning Enthusiast”
- Role
- Lead Data Engineer at Capital One
- Location
- Woodbridge, VA, US
- LinkedIn followers
- 500 followers
About Mohammad Khan
Professional Qualified Data Engineer with around 13+ years of experience in development…
Experience
Lead Data Engineer
Oct 2024 — Present · Leesylvania, VA, US
This project was focused on customer clustering. Used the ETL Data Stage Director to schedule and running the jobs, testing and debugging its components & monitoring performance statistics.Designed and implemented CI/CD pipelines in Azure DevOps to automate deployment of data pipelines, Databricks notebooks, and infrastructure components.Implemented Spark GraphX application to analyze guest behavior for data science segments.Worked on batch processing of data sources using Apache Spark, Elastic search.Developed Big Data solutions focused on pattern matching and predictive modeling.Collaborated with EDW team in, High Level design documents for extract, transform, validate and load ETL process data dictionaries, Metadata descriptions, file layouts and flow diagrams.Designed and maintained modular DBT models following best practices like DRY and modular SQL.A highly immersive Data Science program involving Data Manipulation & Visualization, Web Scraping, Machine Learning, Python programming, SQL, GIT, Unix Commands, NoSQL, MongoDB, Hadoop.Developed Azure Synapse Pipelines to integrate data from multiple structured and unstructured sources, enabling downstream analytics and reporting.Designed and implemented data models for Azure SQL Databases and Data Lakes to support various business use cases, ensuring scalability and performance.Worked on migrating PIG scripts and Map Reduce programs to Spark Data frames API and Spark SQL to improve performance.Created Linked Services for multiple source systems Azure SQL Server, ADLS, BLOB, Rest API.Participated in agile delivery and sprint planning using Azure Boards in conjunction with Git repositories for version control and collaboration.Processing data using Spark (PySpark/Scala) in Azure Databricks.Built scalable ETL pipelines using PySpark in Databricks for batch and streaming data.
Education
National University | Bangladesh
Bachelors in computer science, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.