Anwitha Narsing

Data Engineer @Cencora

Charlotte, NC, US
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Sep 2022 — Present

Data Engineer @Cencora

View department →

NC, US

Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data. Developed Spark scripts by using Scala on AWS as per the requirement. Involved in developing a linear regression model to predict a continuous measurement for improving the observation of wind turbine data developed using spark with Scala API. Hive tables are created as per requirement where Internal or External tables are defined with appropriate static, dynamic partitions and bucketing, intended for efficiency. Extract Real-time feed using Kafka and Spark Streaming and convert it to RDD and process data in the form of Data Frame and save the data as Parquet format in HDFS. Used Spark and Spark-SQL to read the parquet data and create the tables in the hive using the Scala API. Worked in AWS environment for development and deployment of custom Hadoop applications. Experienced in using the spark application master to monitor the spark jobs and capture the logs for the spark jobs. Used Kafka to patch up a customer activity taking after pipeline as a course of action of steady appropriate subscribe supports. Expertise in the deployment of Hadoop, Yarn, Spark, and Storm integration with Cassandra, ignite, and Kafka. Developed spark applications in python (PySpark) on distributed environment to load huge number of CSV files with different schema in to Hive ORC tables. Utilized Agile Scrum Methodology to help manage and organize our team of developers with regular code review sessions Implemented Apache Airflow for authoring, scheduling, and monitoring Data Pipelines. Built an ETL framework for Data Migration from on premise data sources such as Hadoop, Oracle to AWS using Apache Airflow, Apache Sqoop and Apache Spark (PySpark). Transform and analyse the data using PySpark, HIVE, based on ETL mapping. Create Pyspark frame to bring data from DB2 to Amazon S3.

EDUCATION

2014 — 2018

Vardhaman College of Engineering (VCEH)

Bachelor's degree

2022 — 2023

Concordia University-St. Paul

Master's degree

ABOUT ANWITHA NARSING

Profile Overview 7+ years of comprehensive experience in Big Data Technology Stack, Spark Core, Spark SQL, Spark Streaming, Kafka streaming, and Kafka Security and Cloud Services Like Azure. Used ADF and Boomi for Data Ingestion to data lake in Azure ADLS Gen2 from SAP legacy systems. Data computation and analytics used Azure databricks. Used Azure functions for data computations. Worked on Apache Spark performing the Actions, Transformations on RDDs, Data Frames & Datasets Using spark SQL and Spark streaming contexts. Having good experience in spark core, spark SQL and spark streaming. Having good experience in writing Python Lambda functions and calling the API’s. Led the migration of on-premise data warehouses to Azure Synapse Analytics, enhancing scalability, performance, and security, achieving a reduction in operational costs while applying data modeling principles and SQL-based transformations for optimized query performance and improving overall data accessibility. Designed and orchestrated end-to-end data pipelines using Azure Data Factory, Azure Databricks, and Apache Airflow, ensuring efficient data ingestion, transformation, and loading (ETL) processes across various financial data sources. Leveraged SQL, Scala, and PySpark for seamless reporting and analysis, enabling real-time insights into financial data. Developed, tested, and implemented custom ETL solutions using Python, Scala, and PySpark in Azure Data Factory, focusing on complex data transformations and integrating machine learning models for predictive insights into financial systems, improving data modeling and operational efficiency. Managed and optimized large-scale datasets utilizing Azure Blob Storage and HDFS, processing petabyte-scale data with high efficiency and low latency for reliable financial data access while utilizing SQL, Scala, and PySpark for high-speed data processing reducing overall query execution time by optimizing data partitioning strategies and leveraging distributed systems like Apache Flink. Implemented and managed Master Data Management (MDM) solutions using Azure Data Services (Azure Data Factory, Azure SQL Database, Azure Data Lake Storage, and Azure Purview).

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Anwitha Narsing — Data Engineer at Cencora in Charlotte, NC, US | Unifers