Pooja Thakur
Data Engineer | Apache Spark & PySpark | AWS (EMR, S3) | Airflow | Real-Time & Batch Pipelines | Processing Millions of Records Daily | Cassandra | Scalable Data Architectures
- Role
- Data Engineer at Nielsen
- Location
- Nagpur, MH, IN
- LinkedIn followers
- 500 followers
About Pooja Thakur
I am a results-driven Data Engineer passionate about building scalable, high-performance data systems that power business decisions. Currently working at Nielsen, I design and optimize data pipelines that process millions of records daily, enabling reliable analytics and reporting across teams.My expertise lies in transforming complex, heterogeneous data into clean, unified datasets that stakeholders can trust. I work extensively with Apache Spark, PySpark, Airflow, and AWS (EMR, S3, Step Functions) to build resilient batch and real-time data workflows. I also manage hybrid data environments across Cassandra, Sybase, and PostgreSQL, ensuring consistency, synchronization, and low-latency access.What differentiates me is my end-to-end ownership mindset. From architecture design and performance tuning to automated data quality checks and SLA-driven production deployment, I focus on building systems that are not only scalable but reliable and efficient.Before transitioning into data engineering, I worked in infrastructure project management, where I led large-scale projects and coordinated cross-functional teams. That experience strengthened my analytical thinking, stakeholder communication, and structured execution approach, skills I now apply to data engineering projects.I am continuously learning, refining my technical depth in big data and cloud technologies, and looking to contribute to complex, data-driven environments where engineering excellence meets business impact.
Experience
Data Engineer
Jun 2022 — Present · Mumbai, IN
I currently work as a Data Engineer at Nielsen, where I design and maintain high-performance data pipelines that process millions of records every single day. My role sits at the intersection of engineering and analytics, I transform complex, heterogeneous data sources into unified, reliable datasets that power reporting and business decisions.On a typical day, I work with Apache Spark to build scalable processing systems and leverage AWS services to ensure storage and compute can scale seamlessly. I integrate high-velocity Cassandra datasets for real-time, low-latency access and orchestrate workflows using Apache Airflow to make sure stakeholders receive accurate data on time, every time.What makes this role exciting for me is owning the entire data lifecycle, from development to production. I collaborate closely with analytics teams to understand business requirements and translate them into robust, high-performance architectures.Key Contributions & Impact:• I streamlined data ingestion and transformation pipelines, enabling the processing of millions of records daily with improved efficiency.• I implemented automated data quality checks and reconciliation frameworks, significantly improving dataset reliability and trust across teams.• I optimized Spark jobs and queries to enhance scalability and computational performance, reducing processing time and resource consumption.• I managed hybrid data environments across Cassandra and Sybase, ensuring consistency and synchronization.• I led pipelines from design to production while maintaining SLA compliance and minimizing downtime.
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.