Thakur Vaishnavi Singh
Sr Aws Data Engineer @BNY
Signup · Get unlimited contacts
WORK HISTORY
Sr Aws Data Engineer @BNY
NY, US
Develop and add features to existing data analytic applications built with Spark and Hadoop on a Scala, java and Python development platform on the top of AWS services. • Programming using Python, Scala along with Hadoop framework utilizing Cloudera Hadoop Ecosystem projects (HDFS, Spark, Sqoop, Hive, HBase, Oozie, Impala, and Zookeeper etc.). • Involved in developing spark applications using Scala, Python for Data transformations, cleansing as well as validation using Spark API. • Worked on all the Spark APIs, like RDD, Dataframe, Data source and Dataset, to transform the data. • Worked on both batch processing and streaming data Sources. Used Spark streaming and Kafka for the streaming data processing. • Worked on Cloudera distribution and deployed on AWS EC2 Instances. • Developed Spark Streaming script which consumes topics from distributed messaging source Kafka and periodically pushes batch of data to spark for real time processing. • Built data pipelines for reporting, alerting, and data mining. Experienced with table design and data management using HDFS, Hive, Impala, Sqoop, MySQL, and Kafka. • Worked on Apache Nifi to automate the data movement between RDBMS and HDFS. • Created shell scripts to handle various jobs like Map Reduce, Hive, Pig, Spark etc, based on the requirement. • Used Hive techniques like Bucketing, partitioning to create the tables. • Experience on Spark-SQL for processing the large amount of structured data. • Experienced working with source formats, which includes - CSV, JSON, AVRO, JSON, Parquet, etc. • Worked on AWS to aggregate clean files in Amazon S3 and also on Amazon EC2 Clusters to deploy files into Buckets. • Designed and architected solutions to load multipart files which can\'t rely on a scheduled run and must be event driven, leveraging AWS SNS, • Involved in Data Modeling using Star Schema, Snowflake Schema.
EDUCATION
Sree Datta Institute of Engineering & Sciences
Computer Science Engineer
ABOUT THAKUR VAISHNAVI SINGH
9+ years of professional IT experience in BIGDATA using HADOOP framework and Analysis, Design, Development, Documentation, Deployment and Integration using SQL and Big Data technologies as well as Java / J2EE technologies with AWS, AZURE • Experience in Hadoop Ecosystem components like Hive, HDFS, Sqoop, Spark, Kafka, Pig. • Experience in architecting, designing, installation, configuration and management of Apache Hadoop Clusters, MapR, and Horton works & Cloud era Hadoop Distribution. • Good understanding of Hadoop architecture and Hands on experience with Hadoop components such as Resource Manager, Node Manager, Name Node, Data Node and Map Reduce concepts and HDFS Framework. • Expertise in Data Migration, Data Profiling, Data Ingestion, Data Cleansing, Transformation, Data Import, and Data Export through the use of multiple ETL tools such as Informatica Power Centre. • Working knowledge of Spark RDD, Dataframe API, Data set API, Data Source API, Spark SQL and Spark Streaming. • Experience in exporting as well as importing the data using Sqoop between HDFS to Relational Database systems and vice-versa and load into Hive tables, which are partitioned. • Worked on HQL for required data extraction and join operations as required and having good experience in optimizing Hive Queries. • Experience in Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance. • Developed Spark code using Scala, Python and Spark-SQL/Streaming for faster processing of data. • Implemented Spark Streaming jobs in Scala by developing RDD\'s (Resilient Distributed Datasets) and used Pyspark and spark-shell accordingly. • Profound experience in creating real time data streaming solutions using Apache Spark /Spark Streaming, Kafka and Flume. • Experience on Migrating SQL database to Azure Data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and Migrating On premise databases to Azure Data lake store using Azure Data factory. • Good working experience on AWS infrastructure services Amazon Simple Storage Service (Amazon S3), EMR, lambda functions and Amazon Elastic Compute Cloud (Amazon EC2). • Experience in Data Analysis, Data Profiling, Data Integration, Migration, Data governance and Metadata Management, Master Data Management and Configuration Management. • Expertise in designing complex Mappings and have expertise in performance tuning and slowly changing Dimension Tables and Fact tables.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.