Arvind Kale
Senior Data Engineer | Spark • ClickHouse • AWS • Azure | Building High-Performance Data Pipelines | Cost Optimization & Scalable Data Platforms|Blogger | Mentor
- Role
- Data Engineer at Zuno General Insurance
- Location
- Mumbai, MH, IN
- LinkedIn followers
- 500 followers
About Arvind Kale
I design and build scalable data platforms that turn raw data into business intelligence.With experience in Big Data, Cloud, and Data Engineering, I specialize in building high-performance data pipelines using technologies like Spark, Python, ClickHouse, and modern cloud architectures.Recently, I helped re-architect a data pipeline that reduced infrastructure costs by per year while improving performance and query speed.My work focuses on:• Building scalable data pipelines• Optimizing data architecture for cost and performance• Designing analytics platforms for real-time insights• Implementing Spark-based big data processing systemsTech Stack:Python | PySpark | ClickHouse | PostgreSQL | AWS | Azure | Databricks | ETL | Data ModelingI regularly share insights about data engineering, architecture design, and cost optimization.Always open to conversations about:Data Engineering • Big Data Architecture • Cloud Data Platforms
Experience
Data Engineer
Nov 2022 — Present
Experienced Data Engineer with a proven track record in building and optimizing large-scale data pipelines using Azure Data Lake, Databricks, and PySpark. Re-engineered ETL pipelines with PySpark optimization techniques such as caching, partitioning, and broadcast joins, reducing processing time by 20% for daily loads exceeding 1 million records. Automated finance reporting pipelines using PySpark and SQL, achieving an 80% reduction in turnaround time for critical data processes. Designed robust data ingestion frameworks integrating multi-cloud data sources (Azure, AWS, GCP, SAP HANA, Oracle) and handling diverse formats like Parquet, JSON, CSV. Migrated legacy on-premise ETL workloads to cloud-native big data platforms, enhancing system performance and reducing job runtimes by 20%. Optimized performance for massive datasets exceeding 15TB and 12 billion records, implementing advanced tuning strategies for efficient resource utilization. Developed dynamic, event-driven orchestration using Apache Airflow and Python, reducing manual interventions by 80% and improving workflow reliability. Engineered secure, scalable data pipelines for Azure Data Lake, supporting cost-efficient access to infrequent datasets and enabling seamless analytics. Delivered critical production support and SIT/UAT fixes by leveraging PySpark, Spark SQL, and root cause analysis to ensure data pipeline stability.
Education
Trendy Tech
post graduate diploma, Data Engineering
2023 — 2024
SNJC, Maharashtra
Higher secondary education
2013 — 2015
Maharashtra Institute of Technology Aurangabad
Bachelor of Technology - B-Tech, Mechanical Engineering
2015 — 2019
Ramrao Adik Public School, Ahilyanagar Maharashtra
School secondary education
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.