Aayush Joshi
Data Engineer
- Role
- Data Engineer at Heliocampus
- Location
- Washington, DC, US
- LinkedIn followers
- 500 followers
About Aayush Joshi
6 years of technical software development experience and expertise in Big Data, Hadoop Ecosystem, Analytics, Cloud Engineering, Data Warehousing. • Experience in large scale application development using Big Data ecosystem - Hadoop (HDFS, MapReduce, Yarn), Spark, Kafka, Hive, Impala, HBase, Sqoop, Pig, Airflow, Oozie, Zookeeper, Ambari, Flume, Nifi, AWS, Azure, • Sound Involvement with AWS services like Amazon EC2, S3, EMR, Amazon RDS, VPC, Amazon Elastic Load Balancing, IAM, Auto Scaling, Cloud Front, CloudWatch, SNS, SES, SQS, and Lambda to trigger resources. • Experience with building data pipelines using Azure Data Factory, Azure Databricks, and stacking data to Azure Data Lake, Azure SQL Database, Azure SQL Data Warehouse to control and concede database access. • Good experience with Azure services like HDInsight, Stream Analytics, Active Directory, Blob Storage, Cosmos DB, Storage Explorer. • In-profundity understanding/knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, Map Reduce, Spark. • Experience in job workflow scheduling and monitoring tools like Oozie and Zookeeper. • Expertise recorded as a hard copy Hadoop occupation utilizing MapReduce, Apache Crunch, Hive, Pig, and Splunk. • Profound information in developing production-ready Spark applications using Spark Components like Spark SQL, Matplotlib, Graph X, Data Frames, Datasets, Spark-ML and Spark Streaming. • Strong working experience with SQL and NoSQL databases (Cosmos DB, MongoDB, HBase, Cassandra), data modelling, tuning, disaster recovery, backup and creating data pipelines. • Experienced in scripting with Python (PySpark), Java, Scala and Spark-SQL for development, aggregation from various file formats such as XML, JSON, CSV, Avro, Parquet, ORC. • Great experience in data analysis using HiveQL, Hive-ACID tables, Pig Latin queries, custom MapReduce programs and achieved improved performance. • Developing End to End ETL Data pipeline that take the data from surge and loading it into the RDBMS using the Spark. • Experience in configuring Spark Streaming to receive real time data from the Apache Kafka and store the stream data to HDFS and expertise in using spark-SQL with various data sources like JSON, Parquet and Hive. • Experience in ELK stack to develop search engines on unstructured data within NoSQL databases in HDFS. • Created Kibana visualizations and dashboards to view the number of messages processing through the streaming pipeline for the platform.
Experience
Data Engineer
Apr 2024 — Present
Extensively used AWS Athena to import structured data from S3 into other systems such as RedShift to generate reports.• Optimized Hadoop performance by implementing Apache Spark, improving processing speed by 40%.• Created a Data Pipeline utilizing Processor Groups and numerous processors in Apache Nifi for Flat File, RDBMS as part of a Proof of Concept (POC) on Amazon EC2.• Migrated an existing on-premises application to AWS. AWS services such as EC2 and S3 were used for data set processing and storage. Experienced in maintaining a Hadoop cluster on AWS EMR.• Performed end-to-end architecture and implementation evaluations of different AWS services such asAmazon EMR, Redshift, S3, Athena, Glue, and Kinesis. • Developed and implemented ETL pipelines on S3 parquet files in a data lake using AWS Glue.• Developed a cloud formation template in JSON format to utilize content delivery with cross-region replication using Amazon Virtual Private Cloud.• Implemented Columnar Data Storage, Advanced Compression, and Massive Parallel Processing using the Multi-node Redshift technology. • Worked on the code transfer of a quality monitoring application from AWS EC2 to AWS Lambda, as well as the construction of logical datasets to administer quality monitoring on snowflake warehouses.• Migrated a quality monitoring application from AWS EC2 to AWS Lambda, optimizing Snowflake warehouse data processing. • Collaborated with cross-functional teams to gather data requirements and design scalable data architecture, leveraging Airflow for workflow orchestration and coordination. • Ingested data real-time data application of flat files and API’s using Kafka.• Developed data transformation and aggregation workflows using AWS Glue and Apache Spark, enabling real-time and batch processing data.
Education
St. Cloud State University
Master's degree
2020 — 2021
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.