Sai Kumar
AWS Big Data Engineer at Broadridge | Actively Seeking for New Opportunities | Data Analytics| Data Science |SnowFlake |Big Data | SQL | AWS | Hadoop | Azure PySpark | Kafka | Yarn | HDFS | Scala | ETL|hive
- Role
- Data Engineer at Broadridge
- Location
- Charlotte, NC, US
- LinkedIn followers
- 500 followers
About Sai Kumar
Dynamic and motivated IT professional with around 8 years of experience as a Big Data Engineer with expertise in designing data intensive applications using Cloud Data engineering, Data Warehouse, Hadoop Ecosystem, Big Data Analytical, Data Visualization, Reporting, and Data Quality solutions. • Hands on experience across Hadoop Ecosystem that includes extensive experience in Big Data technologies like HDFS, MapReduce, YARN, Apache Cassandra, NoSQL, Spark, Python, Scala, Sqoop, HBase, Hive, Oozie, Impala, Pig, Zookeeper, and Flume. • Built real time data pipelines by developing Kafka producers and Spark streaming applications for consuming. Utilized Flume to analyze log files and write into HDFS. • Experienced with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Dataframe API, Spark Streaming, Pair RDD\'s and worked explicitly on PySpark• Developed framework for converting existing PowerCenter mappings and to PySpark (Python and Spark) Jobs. • Hands on experience in setting up workflow using Apache Airflow and Oozie workflow engine for managing and scheduling Hadoop jobs. • Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR. • Hands-on experience with Amazon EC2, S3, RDS(Aurora), IAM, CloudWatch, SNS, Athena, Glue, Kinesis, Lambda, EMR, Redshift, DynamoDB and other services of the AWS family and in Microsoft Azure. • Proven expertise in deploying major software solutions for various high-end clients meeting the business requirements such as big data Processing, Ingestion, Analytics and Cloud Migration from On-prem to AWS Cloud. • Experience in Work on AWS Databases like Elastic Cache (Memcached & Redis) and NoSQL databases - HBase, Cassandra & MongoDB, database performance tuning & data modeling. • Established connection from Azure to On-premises data center using Azure Express Route for Single and Multi-Subscription. • Created Azure SQL database, performed monitoring and restoring of Azure SQL database. Performed migration of Microsoft SQL server to Azure SQL database. • Experienced in Data Modeling & Data Analysis experience using Dimensional Data Modeling and Relational Data Modeling, Star Schema/Snowflake Modeling, FACT & Dimensions tables, Physical & Logical Data Modeling
Experience
Data Engineer
Aug 2019 — Present · New York, NY, US
Developed Apache presto and Apache drill setups in AWS EMR (Elastic Map Reduce) cluster, to combine multiple databases like MySQL and Hive. This enables to compare results like joins and inserts on various data sources controlling through single platform. • The AWS Lambda functions were written in Scala with cross-functional dependencies that generated custom libraries for delivering the Lambda function in the cloud. • Performed raw data ingestion into S3 from kinesis firehouse, which triggered a lambda function and put refined data into another S3 bucket and wrote to SQS queue as aurora topics. • Writing to the Glue metadata catalog allows us to query the improved data from Athena, resulting in a serverless querying environment. • Created AWS RDS (Relational database services) to work as Hive metastore and could combine 20 EMR cluster\'s meta data into a single RDS, which avoids the data loss even by terminating the EMR. • Built a Full-Service Catalog System which has a full workflow using Elasticsearch, Kinesis, CloudWatch. • Leveraged cloud-provider services when migrating on-prem MySQL clusters to AWS RDS MySQL, provisioned multiple AWS AD forests with AD-integrated DNS, as well as utilized AWS Elasticache for Redis. • Used AWS Code Commit Repository to store their programming logics and script and have them again to their new clusters. • Spin up the EMRs clusters from 30 to 50 nodes which are memory optimized such as R2, R4, X1 and X1e instances with autoscaling feature. • Hive Being the primary query engine of EMR, we have created external table schemas for the data that is being processed.
Education
JNTUH College of Engineering Hyderabad
Bachelor of Technology - BTech, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.