Venkat Ramana
Senior Data Engineer at Ecolab| Hive | Python | Azure | Pyspark | Spark SQL | Azure Databrick| Hadoop | Snow flake| ETL | SQL | Airflow | Agile | Actively looking for new opportunities on C2C/C2H
- Role
- Senior Data Engineer at Ecolab
- Location
- Cleveland, OH, US
- LinkedIn followers
- 500 followers
About Venkat Ramana
Around 9+ years of Experience in Data Engineering, Development, and Implementation as a Data Engineer. • A Good Proficiency experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies. • Strong experience in writing scripts using Python API and Spark API for analyzing the data. • Software professional with commendable experience in maintenance of various web-based application using Java and Big Data Ecosystems and g tools Talend and Informatica. • Well versed in implementing E2E solutions on big data using Hadoop framework. • Proficient in Data Warehousing, Data Mining concepts and ETL transformations from source to target systems. • Worked in multiple Hadoop distributions like Cloudera, Hortonworks, MapR and AWS. • Experience with Developing and Maintaining Applications written for Amazon Simple Storage, AWS Elastic Map Reduce, and AWS Cloud Formation. Imported the data from various sources like AWS S3, local file system into Spark. • Uploaded and processed terabytes of data from various structured and unstructured sources into HDFS using Sqoop. • Very keen in knowing newer techno stack that Google Cloud platform (GCP) adds. • Experience in GCP Dataproc, GCS, Cloud functions, BigQuery. • Implemented POC\'s to migrate map reduce programs into Apache Spark transformations using spark and Scala. • Used Spark-Streaming APIs to perform necessary transformations and actions on the fly. • Experience on Migrating SQL database to Azure Data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and Migrating On premise databases to Azure Data Lake store using Azure Data factory. • Improving performance and optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD\'s. • Hands on experience on developing UDF, DATA Frames and SQL queries in Spark SQL, in Hadoop ecosystem such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and MapReduce programming paradigm. • Experience in using SQOOP for importing and exporting data from RDBMS to HDFS and Hive, performing real time analytics and transfer on data using HBase, HIVE queries & Pig. Extensively used Apache Flume to collect the logs and error messages across the cluster.
Experience
Senior Data Engineer
Jul 2023 — Present · St Paul, MN, US
Designed and developed Sqoop, Linux shell scripts for data ingestion from various data sources of credit Suisse in to HDFS Data Lake. • Created Datasets/ DataFrames from RDDs using reflection and programmatic inference of schema over RDD. • Developed Python Spark programs for processing HDFS Files using RDDs, Pair RDDs, Spark SQL, Spark Streaming, DataFrames, Accumulators, Broadcast variables. • Developed pyspark Kafka streaming programs to integrate various Credit-Suisse source systems to Hadoop, Pyspark programs using various Transformations and operations. • Comprehensive knowledge and experience in process improvement, normalization/de-normalization, data extraction, data cleansing, data manipulation. • Data transformations on HIVE and use static, dynamic partitioning and bucketing for performance improvements. Established and Followed Spark programming best practices. • Performance tuning of Pig, HIVE and Spark jobs used caching / persistence, partitioning and Best practices. • Work with support teams in resolving operational & performance issues. • Research, evaluate and utilize new technologies/tools/frameworks around Hadoop eco system. Involved in all phases of the Software development life cycle (SDLC) using Agile Methodology. • Worked on MySQL database writing queries using Python MySQL connector and MySQL Package. • Implemented CRUD operations for the business logic using RESTFUL services. • Designed and developed use-case, Class, and Object Diagrams for the business requirements. • Used multithreading in programming to improve overall performance. Automated build and deploy process to production environment. Used Jenkins for continuous integration (CI) and continuous deployment (CD). • Used Agile methodologies - Scrums, Sprints, tracking of tasks using JIRA management tool. • Responsible for managing large databases using Panda data frames and MySQL.
Education
Cleveland State University
Master's degree, Information Science/Studies
Jawaharlal Nehru Technological University Kakinada (JNTUK)
Bachelor's Degree, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.