Vaibhav Gaur
Data Engineer @Persistent Systems
Signup · Get unlimited contacts
WORK HISTORY
Data Engineer @Persistent Systems
Ottawa, ON, CA
As a Data Engineer at Persistent Systems, I focus on optimizing data pipelines, ensuring efficient data ingestion, and implementing best practices for large-scale data processing. My responsibilities include: Data Engineering & ETLDesigned and maintained data pipelines to extract, transform, and load (ETL) data from on-prem SQL Server to Snowflake, ensuring data integrity, reconciliation, and zero data loss.Implemented real-time data ingestion using Apache Kafka and PySpark, addressing challenges like schema evolution, partitioning strategies, and late-arriving data.Built secure data pipelines leveraging COA Data Mesh, ensuring smooth data flow from source to destination while maintaining security and compliance. Big Data & Cloud TechnologiesWorked extensively with Databricks, leveraging Delta Lake Time Travel, Zero Copy Clones, Shallow Clones, and Refreshable Clones to optimize data storage and access.Developed and optimized PySpark applications in Azure Databricks, troubleshooting performance bottlenecks by tuning Spark configurations, partitioning strategies, and caching mechanisms.Migrated PySpark notebooks from Databricks to on-prem Windows environments using Databricks REST API & CLI-based execution. Data Governance & SecurityManaged data security protocols when transferring data from on-prem to cloud, ensuring encryption, access controls, and compliance with security policies.Configured role-based access control (RBAC) in Snowflake and Databricks to secure sensitive data. CI/CD & Workflow AutomationIntegrated Azure DevOps, GitLab, and GitHub for CI/CD, enabling smooth deployment of PySpark and SQL-based transformations.Automated pipeline orchestration using Apache Airflow (Cloud-Based) and Tidal (On-Prem Scheduler). API Integrations & ReconciliationImplemented polling mechanisms in Python to track the status of REST API jobs on remote clusters.
EDUCATION
University School of Management Studies, GGSIPU
Bachelor of Technology - BTech
ABOUT VAIBHAV GAUR
With a track record spanning over, I am a dedicated adept at crafting and refining data pipelines to empower actionable insights. My expertise lies in harnessing advanced tools to manage database capacity, fortify security measures, and ensure regulatory compliance. From configuring Snowflake for privacy regulations to designing SSRS report parameters for personalized outcomes, I am committed to optimizing data infrastructure and governance practices:• Data Quality and Governance• Business Intelligence Tools: Power BI, SSRS, SSIS• Azure Services Proficiency: Azure Data Factory, Azure Data Lake Storage, Azure Databricks, Azure Synapse• ETL Process Management: SSIS, Azure Data Factory• Database Expertise: MS SQL Server, Azure SQL, MongoDB, MySQL, PostgreSQL, Snowflake• Project Management: MS Visual Studio, JIRA, SVN, GIT, BITBUCKET• DevOps Integration: Azure Key Vault, Git Version Control• Data Analysis: SQL, T-SQL• Programming Language: Python, Java Let\'s connect to explore synergies in data engineering and analytics! 🤝
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.