Shikhar Rauthan
Data Engineer @ Amazon | AWS Cloud, Python, Apache Spark, Apache Airflow, Apache Kafka, SQL, ETL/ELT Pipeline
- Role
- Data Engineer at Amazon
- Location
- Bengaluru, KA, IN
- LinkedIn followers
- 500 followers
About Shikhar Rauthan
As a Data Engineer at Amazon, I work with the AWS Data Warehouse Team to design and develop scalable and efficient data pipelines, ETL processes, and reporting systems using various AWS services, Python, and Spark. I have reduced internal costs, improved performance, and enhanced data quality and security for multiple projects.Previously, I was instrumental in setting up and managing the Data Warehouse and the Data Lake for Zoomcar, a leading car rental company in India, using Pentaho, Airflow, Redshift, Athena, Lambda, and PySpark. I also created and integrated multiple APIs to pull and push data from third-party advertising vendors, and developed dashboards for business insights on Tableau Server. I have a Master of Computer Applications degree from Vellore Institute of Technology, with a focus on Computer Software Engineering. I am passionate about building data solutions that drive business value and customer satisfaction.
Experience
Data Engineer
Mar 2022 — Present · Bengaluru, IN
AWS Data Warehouse Team· Reduced 1MM+ USD(internal costing) on EMR and saved 30% on SLA by automating job deployment using CDK to deploy CloudFormation templates as inputs to scheduler APIs through CI/CD pipelines and introducing cluster t-shirt sizing.· Designed and developed spark jobs for optimising ETL processes through EMR and Glue for ingesting data from s3 sources to Redshift Data Warehouse and Data Lake. Achieved 20-25% runtime optimisation using partition pruning, spark configs optimisation, filter pushdown, etc.· Saved ~40 PB on storage and 1.4MM+ USD(internal costing) by enabling lifecycle rules on S3 buckets storing archival data.· Worked with AWS legal team and different AWS product teams to enforce the right level of data governance in Datalake.· Migrated Data Warehouse based jobs(Redshift based ETL) to Pyspark based jobs running on EMR. Enhances parquet writes by first caching results in HDFS and then writing to s3 using s3-dist-cp.
Education
Vellore Institute of Technology
Master of Computer Applications - MCA
2018 — 2020
Guru Gobind Singh Indraprastha University
Bachelor of Computer Applications - BCA
2013 — 2016
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.