Naga S.
Senior Data Engineer | AWS | Big Data | PySpark | Cloud ETL | Data Modeling | Hadoop & Spark Ecosystem | DBT & Snowflake Expert
- Role
- Senior Data Engineer at Apple
- Location
- Dallas, TX, US
- LinkedIn followers
- 500 followers
About Naga S.
With over 10 years of experience as a Senior Data Engineer, I have worked across all phases of software development, including requirement analysis, design, development, and maintenance of Big Data applications in the insurance and healthcare domains. My expertise includes AWS cloud technologies such as IAM, EC2, EMR, SNS, RDS, Redshift, Athena, DynamoDB, Lambda, CloudWatch, and S3. I have an in-depth understanding of Hadoop architecture, including HDFS, Name Node, Data Node, MapReduce, and YARN, along with hands-on experience in resource management, task execution, and optimizing large-scale data processing jobs.I have extensive experience in designing and developing Spark applications using Scala, optimizing performance compared to Hive. My work includes integrating Kafka with Spark Streaming for high-speed data processing, developing Spark applications with Spark-SQL for efficient data transformation, and using Cloudera Hadoop YARN for analytics. I have also worked on data ingestion pipelines, collecting log data from web servers and social media using Flume and Kafka, storing it in HDFS for further processing. My expertise extends to data serialization formats such as AVRO, PARQUET, and CSV to handle structured and semi-structured data.My cloud experience includes AWS services like EC2, S3, EMR, CloudFormation, CloudWatch, and Lambda. I have implemented Hive partitioning, dynamic partitioning, and bucketing strategies to improve query performance. Additionally, I have worked with NoSQL databases such as HBase, MongoDB, and Cassandra, and used Sqoop for seamless data movement between HDFS and relational databases. I am also proficient in database technologies like MySQL and Oracle, with expertise in database design, writing complex SQL queries, and developing stored procedures.I have hands-on experience in Azure Data Factory (ADF), designing ETL pipelines for ingesting relational and non-relational data and running Spark jobs in Databricks. My expertise in Snowflake includes bulk data loading from AWS S3 and internal storage using COPY commands, executing import/export operations, and utilizing SnowSQL for efficient data processing. Additionally, I have designed interactive reports and automated dashboards in Tableau, helping organizations track critical KPIs for better decision-making.
Experience
Senior Data Engineer
Jun 2023 — Present · Austin, TX, US
I have extensive experience working with AWS services like Lambda, Glue, and EMR to ingest data from relational and non-relational sources, meeting business requirements efficiently. Utilizing AWS EMR, I have transformed and migrated large datasets between Amazon S3 and databases such as DynamoDB. Additionally, I have leveraged AWS CloudWatch for real-time monitoring and log analysis to ensure system stability. Processed data is stored in Amazon Redshift for further analysis before being loaded into the final database.My expertise includes utilizing Amazon S3 for data storage before and after EMR processing and implementing event logging through AWS EventBridge for better tracking. I have designed and implemented effective data models using DBT, leveraging its transformation and aggregation capabilities. I have also developed and maintained ETL pipelines on AWS using Glue, Lambda, and EMR, streamlining data extraction, aggregation, and consolidation.I have hands-on experience with AWS Glue for data transformation, extraction, merging, and enrichment. Additionally, I have worked on migrating relational database models to the Hadoop ecosystem. Using PySpark, I have built data processing tasks such as reading external data sources, performing enrichment, and loading data into target destinations. My expertise in Spark and Spark-SQL includes reading Parquet data, creating Hive tables, and executing efficient transformations.With strong experience in Linux and RDBMS, I have worked extensively with Sqoop for data ingestion into Hadoop environments. I have managed and reviewed Hadoop and HBase log files, ensuring optimal system performance. Furthermore, I have imported metadata into Hive and migrated existing tables and applications to function within Hive and AWS cloud environments, ensuring efficient data processing and scalability.
Education
University of Central Missouri
Master's degree, Big Data Analytics
Jawaharlal Nehru Technological University, Kakinada
Bachelor of Technology - BTech, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.