Venkata B.
Sr Data Engineer
- Role
- Sr Data Engineer at Citi
- Location
- Dallas, TX, US
- LinkedIn followers
- 500 followers
About Venkata B.
About 8+ years Professional experience in IT including 7+ years of comprehensive experience in Big data Development primarily using Hadoop and Spark Ecosystems.• Good Expertise in ingesting, processing, exporting, analyzing Terabytes of structured and unstructured data on Hadoop clusters in Healthcare, Insurance, and Technology domains.• Expertise in working with Hive optimization techniques like Partitioning, Bucketing, vectorizations and Map side-joins, Bucket-Map Join, skew joins, and creating Indexes.• Hands on experience with AWS (Amazon Web Services), Elastic Map Reduce (EMR), Storage S3, EC2 instances and Data Warehousing.• Good Working Knowledge on working with AWS cloud services like EMR, S3, Redshift, EMR cloud watch, for big data development.• Launched EC2 instances with various AMI’s. Integrated EC2 instances with various AWS tool by using IAM roles. Created Images of critical EC2 instances and used those images to spin up a new instance in different AZ’s.• Experience in Google API’s geocode, translate etc. and well versed with designing best algorithms to use paid API’s efficiently and profitably. • Experience in creating dashboards in Stack driver. Can setup alerting and create custom metrics using google API developer tools.• Experience in developing enterprise level solution using batch processing (Using apache Pig) and streaming framework (Using Spark Streaming, apache Kafka & Apache).• Experienced in handling various file formats like AVRO, Parquet, ASCII, XML, JSON.• Experience in gathering requirements, analyzing requirements, providing estimates, implementation, and peer code reviews.• Have good exposure with the star, snowflake schema, data modelling and work with different data warehouse projects.• Expertise in writing DDLs and DMLs scripts in SQL and HQL for analytics applications in RDBMS and Hive.• Hands on experience in setting up workflow using Airflow and Oozie workflow for managing and scheduling Hadoop jobs.• Hands on experience building streaming applications using Spark Streaming and Kafka with minimal/no data loss and duplicates.• Skilled on streaming data using Apache Spark, migrating the data from Oracle to Hadoop HDFS using Sqoop.• Experience in importing and exporting data from HDFS to RDBMS systems like Teradata (Sales Data Warehouse), SQL-Server, and Non-Relational Systems like HBase using Sqoop by efficient column mappings and maintaining the uniformity.• Experience in various python modules as NumPy, pandas, keras and tensor flow for Machine learning.
Experience
Sr Data Engineer
Aug 2020 — Present
Developed Spark application for loading CSV file data and applying business validation on data frame to find invalid and valid data frames. Wrote a valid data frame into the actual Hive partition table and invalid data frame into error table, partitioned by load date and load type.• Working currently as a Data Engineer in Credit Risk Scoring Application for many portfolios in GCB (Global Consumer Banking).• Hands on experience in machine learning, big data, data visualization, Python development, Java, Linux, SQL, GIT/GitHub• Worked with google data catalog and other google cloud APIs for monitoring, query and billing related analysis for BigQuery usage.• Experience in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns• Performed several analytics on the Data Lake using PySpark.• Optimizing of existing algorithms in Hadoop using Spark Context, Spark -SQL, Data Frames and RDD’s.• Worked on several mechanisms for getting the data like Data Ingestion, Data Standardization, Data Quality, Featuring engineering and Scoring.• Modeled Hive partitions extensively for data separation and faster data processing and followed Hive best practices for tuning.• Interacting with Financial services stake holders about the strategies to ensure quality of high priority by building various analytical dashboards.• Extensively worked on PySpark programming for automation and connecting to different ecosystems.• Involved in the development of real time streaming applications using PySpark, Apache Flink, Kafka, Hive on Distributed Hadoop Custer.
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.