Prakash Mirthipati

Senior Data Engineer | Spark, Redshift, AWS, Oracle | Building Petabyte-Scale Data Products & Reliable Distributed Pipelines | Zero-to-One Product & Data Engineering Leader

Role
Data Engineer Iii at Amazon
Location
Seattle, WA, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Prakash Mirthipati

I work at the intersection of data engineering, distributed systems, and product thinking-building reliable data platforms that power analytics and machine learning.I design AWS data systems processing petabyte-scale datasets and billions of records daily across Spark and Redshift pipelines.Earlier in my career, I built an affiliate analytics platform from scratch for the world’s largest online gaming company at the time, replacing a manual Excel-based system. That experience taught me to love solving data problems at any size.My work centers on large-scale AWS data platforms using technologies such as Spark, Redshift, Step Functions, Kinesis, Glue, and S3. These systems support high-performance batch and near-real-time pipelines, enabling analytics, experimentation, and machine learning across large organizations.My experience spans the evolution of modern data architecture-from OLTP systems to enterprise data warehouses, data marts, operational data stores, data lakes, and modern lakehouse platforms. I’ve worked extensively with MPP systems including Redshift, Teradata, and Oracle, along with distributed processing frameworks such as Spark and Hive.I’ve owned Tier-1 datasets serving hundreds of downstream consumers, where durable schemas, automated validation, and strong quality gates are essential to maintaining trust at scale. The pipelines behind these datasets ingest tens of terabytes of data daily and produce analytics-ready outputs that inform product, operational, and business decisions.A strong focus on pipeline optimization, infrastructure efficiency, and observability has improved end-to-end runtimes by 20–30% while reducing operational costs. I translate complex business questions into scalable, self-service data products that enable teams to confidently use data for analytics, machine learning, and experimentation.Core areas: Distributed Data Processing • Data Platform Architecture • Large-Scale ETL/ELT • Lakehouse & Data Warehouse Systems • AWS Data Infrastructure

Experience

  1. Data Engineer Iii

    Amazon

    Nov 2023 — Present · Bellevue, WA, US

    Managed high-scale analytical data products processing ~2TB of data daily (peaks of 200M events/hour) for analytics and machine learning use cases.Built and operated scalable AWS-based batch and near real-time data pipelines delivering reliable, ML-ready datasets to a wide range of downstream consumers.• Designed adaptable, high-performance data models in Amazon Redshift to support evolving analytical requirements while ensuring long-term maintainability• Developed and standardized Amazon S3-based ingestion pipelines (JSON extraction, flattening, Parquet conversion) for consistent processing of semi-structured data• Processed high-volume event streams (110M/hour avg, 200M/hour peak), ensuring data availability and consistency for critical business use cases• Partnered with data scientists and analytics stakeholders to deliver analysis-ready and ML-ready datasets, reducing time-to-insight and downstream rework• Improved end-to-end pipeline runtimes by ~20–30% through Spark SQL optimization, Redshift schema design, and workload management tuning• Implemented data quality validation, monitoring, and operational workflows for Tier-1 datasets, ensuring reliability during peak events (e.g, Prime Day, holiday traffic)• Migrated legacy relational workloads to cloud-native AWS platforms, improving system reliability and operational resilience• Managed migration from Redshift-based pipelines to Spark-based architectures, enabling scalable batch processing and improved ML dataset delivery while maintaining SLAs• Reduced infrastructure costs by ~$250K through performance optimization and efficient resource utilization

Education

  • Jawaharlal Nehru Technological University

    Bachelor of Technology, Computer Science and Engineering

  • Dr. B.R.A.G.M.R.Polytechnic, Rajahmundry

    Diploma, Computer Engineering

Skills

  • Erwin
  • Agile Methodologies
  • Oracle
  • Pl/Sql
  • Requirements Analysis
  • Basel Iii
  • Hive
  • Apache Spark
  • Informatica
  • Control-M
  • Unix Shell Scripting
  • Software Development Life Cycle (Sdlc)
  • Core Java
  • Shell Scripting
  • Javascript
  • Oracle Developer 2000
  • Software Project Management
  • Business Objects
  • Amazon Redshift
  • Basel Ii
  • Sdlc
  • Sql
  • Tableau
  • Liquidity Reporting
  • Online Affiliates
  • Jquery
  • Scala

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Prakash Mirthipati — Data Engineer Iii at Amazon in Seattle, WA, US | Unifers