Rajeshkumar Sahu
Sr Data Engineer at Comcast
- Role
- Sr Data Engineer at Comcast
- Location
- Mount Laurel, NJ, US
- LinkedIn followers
- 500 followers
About Rajeshkumar Sahu
Over 15+ Years of extensive IT experience primarily in data world building large scale and also real time data warehousing solution using various platforms/tools like using Spark,AWS,Databricks,Minio, Kubernetes/docker, Hadoop BigData, and Licensed ETL tools like Talend, Informatica, DataStage etc• 4+ years of experience working in most of the AWS components like S3, Lambda, SQS, SNS, Athena, Glue, Redshift, kinesis etc. using Databricks.• 2+ years of experience working with Kubernetes & dockers using spark with scala/python and minio as on-prem storage.• 3+ years of comprehensive experience as a BigData engineer using various Hadoop ecosystem components such as hive, sqoop, Oozie, Spark, kafka etc.• Experience working with Databricks notebook to develop spark solution using (scala/python) to read/write to AWS/Minio platforms.• Expertise in reading real time data from Splunk and kafka, Kinesis, Elasticdb etc and load it into target data layer.• Has very strong historical background working with ETL tools like Informatica PowerCenter, DataStage, Talend etc with databases like Oracle, SQL Server, DB2, Teradata etc.• Strong experience in writing Unix shell scripts, Python scripts, automating ETL/bigdata jobs, error handling, and auditing mechanisms. • Have good hands-on experience working with scheduling tools like Apache Airflow, Autosys, Atomic UC4 and Cron schedular.• Have good understanding traditional Data warehousing designs and concept.KEY Achievements:• Optimized data and compute costs by implementing lifecycle policies version control, avoiding cross region vpc endpoint connection, tuning the spark code, retiring unwanted processes there by reducing Databricks/AWS cost by 50% over all making this as significant achievement• Developed a reusable ingestion framework using spark with scala to read data from various sources like AWS, RDBMS, Minio, Data Lake and load data into our native S3 layer, this has reduced more than 20% development efforts on all our ETL pipeline development.• Developed source file automation framework in airflow which will send auto email to source team about missing file and the pipeline will resume automatically once file arrived reducing 20% operational cost due to manual intervention.• Developed a reusable DQ framework to validate the data accuracy reducing significant development time and code redundancy.• Developed a EMR serverless data processing framework reducing 25 % operational cost reducing compared to Databricks.
Experience
Sr Data Engineer
Jun 2019 — Present · Philadelphia, PA, US
Hands on experience developing ETL Pipeline using multiple different tools and components like S3, Lambda, SQS, SNS, Glue, CloudWatch, EMR Serverless, Databricks, Spark-scala, Kinesis, Athena, RedShift, Data Lake etc.• Currently leading the Operational team by supporting 250+ ETL pipelines built using various component running over Databricks and scheduled in Airflow.• Working with business stake holders, understanding the Intake requirement, working with scrum master to create appropriate User Story and working with the developer until delivered.• Working with architecture calls to review the end-to-end process and provide valuable inputs wherever necessary.• Worked closely with platform team to optimize data and compute costs by implementing lifecycle policies version control, avoiding cross region vpc endpoint connection, tuning the spark code, retiring unwanted processes there by reducing Databricks/AWS cost by 50% over all making this as significant achievement• Lead the efforts to develop some reusable and automation framework which significantly reduced the manual intervention and development time.• Streamlined the legacy process and optimized it to reduce cost and achieve operational excellency.• Implemented CI/CD using github actions, Opentofu and code scanning tools like checkmarks, Synk and SonarCube to reduce the code vulnerability and adhere to company’s security standards.• Developed a reusable ingestion framework using spark with scala to read data from various sources like AWS, RDBMS, Minio, Data Lake and load data into our native S3 layer, this has reduced more than 20% development efforts on all our ETL pipeline development.• Developed source file automation framework in airflow which will send auto email to source team about missing file and the pipeline will resume automatically once file arrived reducing 20% operational cost due to manual intervention.
Education
Mumbai University
BE, Electronics
2004 — 2008
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.