Kiran Hake
Data Engineer@TCS |Hadoop|Pyspark|MySQL| |Oozie| ETL|Delta lake|Databricks|ADF|Hive| Jira| | Sqoop |impala|
- Role
- Big Data Engineer at Tata Consultancy Services
- Location
- Pune Division, MH, IN
- LinkedIn followers
- 500 followers
About Kiran Hake
Data Engineer with 3.10 years of experience in Big Data technologies, specializing in designing and building scalable data pipelines using the Spark framework (PySpark & Spark SQL). Strong expertise in data ingestion, transformation, and processing of large-scale structured and semi-structured datasets across distributed environments. Hands-on experience working with Hadoop ecosystem components including HDFS, Hive, and Impala, with a solid understanding of data warehousing concepts and ETL/ELT methodologies. Proficient in writing complex SQL queries, performance tuning, and optimizing Spark jobs for improved execution efficiency and reduced latency. Experienced in developing batch data pipelines, automating data quality validation checks, and implementing robust data processing workflows. Skilled in Python for data manipulation, scripting, and pipeline orchestration. Familiar with partitioning, bucketing, and file formats such as Parquet and ORC for efficient storage and querying.
Experience
Big Data Engineer
Mar 2023 — Present · Pune District, IN
Currently working on building a scalable and high-performance data engineering solution for a banking (BFSI) client, handling large volumes of daily transactional and operational data from multiple heterogeneous sources including RDBMS, flat files, and external systems.Designing and developing robust, fault-tolerant end-to-end data pipelines for data ingestion, processing, and storage using Apache Spark (PySpark) and the Hadoop ecosystem. Implementing efficient data ingestion frameworks using Sqoop and Spark to load data into HDFS, followed by scalable ETL/ELT transformations using Spark SQL and Hive.Optimizing large-scale data processing by applying advanced techniques such as partitioning, bucketing, predicate pushdown, broadcast joins, and caching to improve performance and reduce execution time. Designing Hive tables using columnar storage formats like Parquet and ORC for efficient querying and storage optimization.Actively working on workflow orchestration and scheduling using Oozie, ensuring reliable job execution and dependency management. Implementing automated data quality validation frameworks, data reconciliation, and error-handling mechanisms to maintain high data accuracy and consistency.Collaborating with cross-functional teams in an Agile environment using JIRA, ensuring timely delivery of scalable, secure, and high-quality data solutions aligned with business requirements.
Education
Savitribai Phule Pune University
Bachelor of Engineering - BE, Computer Engineering
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.