Sujith Kumaravel
Data Engineer @CGI
Signup · Get unlimited contacts
WORK HISTORY
Data Engineer @CGI
Montreal, QC, CA
Designed and deployed scalable batch and streaming ETL pipelines using AWS Glue (PySpark) and Kinesis Data Streams, enabling near real-time data processing for critical business applications.Built metadata-driven ELT frameworks with AWS Step Functions and Lambda, reducing hard-coded logic and increasing reusability across datasets.Orchestrated cross-system workflows using Apache Airflow (Amazon MWAA) to integrate Redshift, S3, and external APIs with robust dependency handling.Developed high-quality data validation layers (schema checks, anomaly detection) in PySpark, improving data reliability and reducing downstream errors by 40%.Modeled and optimized dimensional schemas in Amazon Redshift, including fact/dimension/bridge tables, for enterprise BI and dashboarding.Integrated third-party platforms such as Salesforce and SAP into the AWS data lake using custom Python connectors and ingestion jobs.Led cloud-native containerization of transformation workloads via Docker and ECS, supporting isolated and fault-tolerant execution.Implemented CI/CD pipelines using Jenkins, GitHub Actions, and AWS CodeBuild for automated testing and deployment of data assets.Delivered executive-ready Power BI dashboards sourced from Redshift and Athena, tracking business SLAs and operational KPIs.Automated infrastructure provisioning using Terraform, ensuring consistent dev-to-prod environments across AWS services.Configured CloudWatch and SNS alerting mechanisms for real-time failure notifications and SLA tracking.Conducted training sessions on AWS-native tools to upskill team members and streamline platform adoption.
EDUCATION
BSA Crescent Institute of Science and Technology
Bachelor of Technology - BTech
Concordia University
Master of Applied Computer Science
ABOUT SUJITH KUMARAVEL
I’m a passionate Data Engineer and BI Developer with 4+ years of hands-on experience designing, building, and optimizing data ecosystems across Azure and AWS. From real-time ETL pipelines to modern data lakehouses, I specialize in creating scalable, resilient, and production-grade data solutions that drive business insight and impact.I’ve worked extensively with tools like Azure Data Factory, Databricks (PySpark), Snowflake, Apache Airflow, and AWS Glue to transform raw data into reliable, analytics-ready datasets. My work spans across orchestrating pipelines, architecting data models, securing infrastructure, and building CI/CD workflows for seamless deployments.Beyond engineering, I actively collaborate with analysts, data scientists, and business stakeholders to enable predictive modeling, ML workflows, and dashboarding via Power BI and Tableau. I’m also a strong advocate for data governance, observability, and automation—ensuring quality, compliance, and operational excellence. What I bring:Expert in ETL/ELT frameworks, cloud warehousing, and PySpark.Skilled in both batch and streaming data processing.Passionate about clean code, reusable frameworks, and DevOps best practices.Proven experience leading data platform migrations, mentoring junior engineers, and delivering insights that scale.Let’s connect if you’re working on data modernization, building analytics platforms, or looking for someone to bring structure to chaos in your data landscape.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.