Vivek B

Senior Data Engineer @Silicon Valley Bank

Dallas-Fort Worth, TX, US
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Sep 2020 — Present

Senior Data Engineer @Silicon Valley Bank

View department →

CA, US

Developing ETL pipelines to move on-prem data (data sources that include Flat Files, Mainframe Files, and Databases) to AWS S3 using Talend, PySpark Created and embedded python modules in the ETL pipeline to automatically migrate data from S3 to Redshift using AWS Glue. • Built bi-directional ingestion framework from various cloud endpoints like AWS S3 buckets. • Responsible for implementing a generic framework to handle different data collection methodologies from the client primary data sources, validate transform using spark and load into S3 • Used AWS cloud product suites (S3, EMR, Glue, Lambda, SQS, Redshift), Hive, Hadoop, Spark SQL • Operated on AWS cloud S3, EC2, IAM policies, SQS, Lambda, AWS Sage maker • Designed solution to handle AWS Glue Timeout issue, automated reruns, and notification process. • Designed and architected solutions to load multipart files which can\'t rely on a scheduled run and must be event driven, leveraging AWS SNS, SQS, Lambda and Glue. • Created Lambda function to run the AWS Glue job based on the dened Amazon S3 event. • Automated tasks of extracting metadata and lineage from tools using Python scripts and saved 70+ hours’ manual efforts. • Created program in python to handle PL/SQL functions like cursors and loops which are not supported by snowflake. • Worked on the Spark SQL and Spark Streaming modules of Spark and used Scala and Python to write code for all Spark use cases. • Exploring with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark-Context, Spark-SQL, Data Frame and Pair RDD\'s. • ETL pipelines in and out of data warehouse using combination of Python and Snowflakes SnowSQL Writing SQL queries against Snowflake. • Developed data warehouse model in snowflake for over 100 datasets using whereScape. • Creating Reports in Looker based on Snowflake Connections • Validating the data from SQL Server to Snowflake to make sure it has Apple to Apple match.

EDUCATION

2008 — 2012

Jawaharlal Nehru Technological University

Bachelor of Technology - BTech, Computer Science

ABOUT VIVEK B

8+ years of IT experience in Architecture, Analysis, design, development, Testing, implementation, maintenance and support with experience in developing strategic methods for deploying BIG DATA technologies to efficiently solve Big Data processing requirement on Cloud/On Premise. • 4 years of Experience on BIG DATA using HADOOP framework and related technologies such as HDFS, HBASE, Map Reduce, HIVE, PIG, FLUME, OOZIE, POSTGRES, SQOOP, TALEND, IMPALA and ZOOKEEPER,AWS. • Around 2 years of experience on Apache SPARK STORM and KAFKA. • Experience in AWS infrastructure usingEC2, Auto-Scaling in launching EC2 instances, Elastic Load Balancer, Elastic Beanstalk, S3, Glacier, Cloud Front, RDS, VPC, Route53, Cloud Watch, Cloud Formation, IAM, SNS,EBS. • Working knowledge on GCP tools like Big Query, Pub/Sub, Cloud SQL and Cloud functions. • Build data pipelines in airflow in GCP for ETL related jobs using different airflow operators. • Developed end to end job automations using Apache Airflow using Python DAG’s. • Worked on data load from various sources i.e, Oracle, MySQL, DB2, MS SQL Server, Cassandra, Nifi, MongoDB, Hadoop using Sqoop and PYTHON SCRIPT. • Experience with Snowflake cloud data warehouse and AWS S3 bucket for integrating data from multiple source system which include loading nested JSON formatted data into snowflake table. • Experienced with SPARK SREAMING API to ingest data into SPARK ENGINE from KAFKA. • Experience on Migrating SQL database to Azure data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and Migrating On premise databases to Azure Data lake store using Azure Data factory. • Excellent understanding and knowledge of NOSQL database HBASE and CASSANDRA. • Extensive experience with Amazon Web Services (AWS), Open stack, Docker, AWS Cloud Formation, AWS Cloud Front and Experience in using containers like Docker and familiar with Jenkins. • Designed and architected solutions to load multipart files which can\'t rely on a scheduled run and must be event driven, leveraging AWS SNS, SQS, Lambda and Glue. • Worked extensively with Dimensional MODELING (OLAP), DATA MIGRATION, DATA CLEANSING, DATA PROFILING, and ETL Processes features for data warehouses. • Experience in all stages of SDLC (Agile, Waterfall), writing Technical Design document, Development, Testing and Implementation of Enterprise level Data mart and Data warehouses.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Vivek B — Senior Data Engineer at Silicon Valley Bank in Dallas-Fort Worth, TX, US | Unifers