Unlimited Email and Phone number

Everything you need to reach more inboxes and close more deals

Unlimited Contact Only at$79

Sign up now

Naveen P.

Data Engineer @Capital One

Plano, TX, US
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Sep 2024 — Present

Data Engineer @Capital One

Designed and optimized scalable ETL pipelines using PySpark and Spark Streaming on AWS EMR, processing terabytes of structured and semi-structured telecom data.Developed and managed real-time ingestion frameworks using Apache NiFi and AWS Lambda, enabling low-latency data delivery to analytical systems.Implemented data lakehouse architecture on Amazon S3 with Apache Iceberg, improving query performance and schema evolution handling.Integrated Aurora PostgreSQL and Snowflake for downstream analytical consumption, building efficient ELT workflows and optimizing T-SQL queries.Built orchestration workflows in Apache Airflow to schedule, monitor, and recover critical production pipelines with minimal downtime.Deployed containerized Spark applications on Amazon EKS and automated CI/CD pipelines using Docker and AWS CodePipeline for consistent delivery.Configured AWS CloudWatch and custom alerts for EMR, NiFi, and Airflow to monitor pipeline performance and ensure SLA compliance.Implemented fine-grained access controls using AWS Lake Formation, improving data governance and security posture.Created and tuned Python-based automation scripts for metadata-driven ingestion and data quality checks, reducing manual intervention by 40%.Partnered with telecom domain experts and analysts to provision clean, curated data for advanced AI/ML analytics use cases.Conducted peer code reviews, best practice workshops, and knowledge sharing sessions focused on PySpark, AWS Data Services, and data security standards.

EDUCATION

N/A

University of Central Missouri

Master of Science - MS, Computer Science

ABOUT NAVEEN P.

As a data engineer with over 8+ years of hands-on experience, currently at Capital One,I specialize in designing and optimizing scalable ETL pipelines using PySpark and Azure Databricks. My work involves processing and transforming multi-terabyte datasets, building real-time data ingestion systems with Kafka, and implementing Azure-based data solutions to enhance cost efficiency and infrastructure performance. My technical expertise extends to Snowflake, MongoDB, and ElasticSearch, enabling fast and reliable data access for business analytics. Previously, I contributed to healthcare data pipeline migration projects at The Cigna Group, leveraging Azure Databricks, AWS S3, Glue, and Redshift for seamless cloud transitions. My mission is to empower organizations by delivering robust, high-performance data solutions that facilitate actionable insights. Motivated by innovation and collaboration, I aim to create efficient, future-ready data platforms that align with organizational goals.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.