Lakshmi V
Sr. Data Engineer at NationWide
- Role
- Senior Data Engineer at Nationwide
- Location
- Columbus, OH, US
- LinkedIn followers
- 500 followers
About Lakshmi V
Over 8+ years of experience in Designing, Developing, and integrating applications using Hadoop, Hive, PIG, in all three Bigdata platforms Cloudera, Hortonworks, MapR, Snowflake, Apache Airflow in which 5 years of experience on Data Engineering and 4years of experience on Data Warehouse. • Experience writing pig and hive scripts. • Experience in writing Map Reduce programs using Apache Hadoop for analyzing Big Data. • Proficient in the Integration of various data sources with multiple relational databases like Oracle11g /Oracle10g/9i, Sybase12.5, Teradata and Flat Files into the staging area, Data Warehouse and Data Mart. • Expertise in Tuning & Optimizing the DB relevant issues (SQL Tuning). • Hands on experience in writing Ad-hoc Queries for moving data from HDFS to HIVE and analyzing the data using HIVE QL. • In depth knowledge of Hadoop Architecture and Hadoop daemons such as Name Node, Secondary Name Node, Data Node, Job Tracker and Task Tracker. • Good working knowledge on Snowflake and Teradata databases. • Extensively worked on Spark using Scala on cluster for computational (analytics), installed it on top of Hadoop performed advanced analytical application by making use of Spark with Hive and SQL/Oracle/Snowflake. • Expertized in Python data extraction and data manipulation, and widely used python libraries like NumPy, Pandas, and Matplotlib for data analysis. • Extensively worked on other machine learning libraries such as Seaborn, Scikit learn for machine learning and familiar working with TensorFlow, NLTK for deep learning • Experienced in creating SnowFlake Multi-cluster Size and Credit Usage. • Played key role in Migrating Teradata objects into SnowFlake environment. • Experience with Snowflake Multi-Cluster Warehouses • Experience with Snowflake Virtual Warehouses. • Having In-depth knowledge of Data Sharing in Snowflake. • Have a knowledge of Snowflake Database, Schema and Table structures. • Experience in using Snowflake Clone and Time Travel. • Solid understanding and experience with extract, transform, load • Implementing data movement from File system to Azure Blob storage using python API • Implemented PoC for running ML models on Azure ML studio • Written Kafka consumer Topic to move data from adobe clickstream Json object to Datalake.• Experience on working with file structures such as text, sequence, parquet and Avro file formats. • Implemented Time Series Map Reduce paradigm using Java and Spark. • Expertise in using Sqoop & Spark to load data from MySQL/Oracle to HDFS or HBase.
Experience
Senior Data Engineer
Aug 2020 — Present · Columbus, OH, US
Developed Spark applications using Scala. • Performance analysis of batch jobs by using Spark Tuning parameters. • Enhanced and optimized Spark/Scala/ pyspark jobs to aggregate, group and run data mining tasks using the Spark framework. • Worked on aws tools like Kinesis, DynamoDB, S3. • Importing and exporting data into HDFS and hive using Sqoop and Kafka with batch and streaming. • Worked on MySQL RDBMs db. as backend database to store monitoring information about CCPA project. • Used Service now platform to open and close tickets of CCPA project. • Involved in complete Big Data flow of the application data ingestion from upstream to HDFS, processing the data in HDFS and analyzing the data using several tools. • Imported the data from various formats like JSON, ORC and Parquet to HDFS cluster with compressed for optimization. • Experienced on ingesting data from RDBMS sources like - Oracle, SQL Server and Teradata into HDFS using Sqoop. • Deployed Pyspark applications and developed in Databricks cluster. • Experience in managing and reviewing huge Hadoop log files. • Importing and exporting data into HDFS and hive using Sqoop and Kafka with batch and streaming. • Experienced with Spark-Streaming APIs to perform transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real time and Persists into HBase. • Performance analysis of Spark streaming and batch jobs by using Spark tuning parameters. • Enhanced and optimized Spark/python jobs to aggregate, group and run data mining tasks using the Spark framework. • Installed, configured and developed various pipeline activities with Nifi using various processors such as Sqoop processor, Kafka processor, HDFS Processor, File Processors etc. • Created Data Pipelines as per the business requirements and scheduled it using Oozie. • Used Hive to join multiple tables of a source system and load them to Elastic search tables.
Education
Vijay Maruti
Bachelor's degree, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.