Hemanth R.
Data engineer | Building Scalable Data Platforms | AWS | Airflow | python |SQL | Spark
- Role
- Data Engineer at Nationwide
- Location
- Charlotte, NC, US
- LinkedIn followers
- 500 followers
About Hemanth R.
5 years of professional Data Engineer experience with 5 years of expertise in Big Data, Hadoop Ecosystem, Cloud Engineering, Data Warehousing.• Sound Experience with AWS services like Amazon EC2, S3, EMR, Amazon RDS, VPC, Amazon Elastic Load Balancing, IAM, Auto Scaling, Cloud Front, CloudWatch, and Lambda to trigger resources.• Experience in building data pipelines using Azure Data Factory, Azure Databricks, and loading data to Azure Data Lake, Azure SQL Database, Azure SQL Data Warehouse to control and grant database access.• Good experience with Azure services like HDInsight, Stream Analytics, Active Directory, Blob Storage, Cosmos DB, Storage Explorer. • Strong Hadoop and platform support experience with all the entire suite of tools and services in major Hadoop Distributions – Cloudera, Amazon EMR, Azure HDInsight, and Hortonworks. • Proficient in handling and ingesting terabytes of Streaming data (Kafka, Spark streaming, Strom), Batch Data, Automation and Scheduling (Oozie, Airflow).• Profound knowledge in developing production-ready Spark applications using Spark Components like Spark SQL, DataFrames, Datasets, Spark-ML and Spark Streaming.• Expertise in developing multiple confluent Kafka Producers and Consumers to meet business requirements. Store the stream data to HDFS and process it using Spark.• Strong working experience with SQL and NoSQL databases (Cosmos DB, MongoDB, HBase, Cassandra), data modeling, tuning, disaster recovery, backup and creating data pipelines.• Experienced in scripting with Python (PySpark), Scala and Spark-SQL for development, aggregation from various file formats such as XML, JSON, CSV, Parquet.• Great experience in data analysis using HiveQL, Hive-ACID tables, Pig Latin queries, custom MapReduce programs and achieved improved performance.• Extensive knowledge in all phases of Data Acquisition, Data Warehousing (gathering requirements, design, development, implementation, testing, and documentation), Data Modeling (analysis using Star Schema and Snowflake for FACT and Dimensions Tables), Data Processing and Data Transformations (Mapping, Cleansing, Monitoring, Debugging, Performance Tuning and Troubleshooting Hadoop clusters).• Experience in monitoring document growth and estimating storage size for large MongoDB clusters as part of the data life cycle management.• Hands-on experience on Ad-hoc queries, Indexing, Replication, Load balancing, Aggregation in MongoDB.
Experience
Data Engineer
Aug 2022 — Present · Charlotte, NC, US
Involved in developing batch processing applications that require functional pipelining using Spark APIs.• Involved in building a data pipeline and performed analytics using AWS stack (EMR, EC2, S3, RDS, Lambda, Glue, Redshift).• Collaborated with client team to transform data and integrate algorithms and models into automated processes.• Utilized Spark’s in memory capabilities to handle large datasets on S3 Data Lake. Loaded data into S3 buckets, then filtered and loaded into Hive external tables.• Strong Hands-on experience in creating and modifying SQL stored procedures, functions, views, indexes, and triggers.• Performed ETL operations using Python, Spark SQL, S3 and Redshift on terabytes of data to obtain customer insights.• Used programming skills in Python to build robust data pipelines and dynamic systems.• Good Understanding of other AWS services like S3, EC2 IAM, RDS Experience with Orchestration and Data Pipeline like AWS Step functions/Data Pipeline/Glue.• Integrated data from a variety of sources, assuring that they adhere to data quality and accessibility standards.• Experience in building data transformation and processing solutions.• Has strong knowledge of large-scale search applications and building high volume data pipelines.• Experience in Writing ETL (Extract / Transform / Load) processes, designs database systems and develops tools for real-time and offline analytic processing.
Education
University at Albany
Masters in Applied Mathematics, Data science
2020 — 2021
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.