Sahitha Rao

Senior Data Engineer @New York State Education Department

McKinney, TX, US
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Jan 2024 — Present

Senior Data Engineer @New York State Education Department

View department →

IL, US

Architected enterprise-scale data ingestion pipelines using Azure Data Factory and Databricks for ingesting structured and unstructured education data.• Built Databricks clusters with auto-scaling enabled, optimizing for cost and performance across batch and interactive workloads.• Designed Delta Lake architecture on Azure Data Lake Storage (ADLS) Gen2, ensuring data quality, versioning, and efficient data governance.• Implemented real-time streaming pipelines using Azure Event Hubs and Databricks Structured Streaming for operational analytics and dashboards.• Developed PySpark and Spark SQL transformations that handle billions of education records while maintaining sub-minute latency for critical workloads.• Implemented end-to-end CI/CD pipelines using Azure DevOps, automating testing and deployment of Databricks notebooks and Delta tables.• Set up comprehensive monitoring and alerting using Azure Monitor, tracking pipeline health, performance metrics, and data quality SLAs.• Managed data security with Azure Key Vault, Databricks ACLs, and column-level encryption to comply with FERPA and education data privacy regulations.• Built data quality frameworks in Databricks using custom PySpark functions and Great Expectations for automated validation and anomaly detection.• Designed dimensional models in Databricks SQL, creating fact and dimension tables that enable fast analytics queries for state-level reporting.• Orchestrated complex step Databricks workflows using Azure Data Factory, managing dependencies and error handling gracefully.• Optimized Databricks SQL queries reducing execution time by 70% through partitioning, caching, and Photon acceleration.• Deep experience with CMS Stars Reporting - transforming raw claims into HEDIS measures that actually impact 5-star ratings• Created Delta Live Tables pipelines processing millions of Medicare Advantage member months for timely Stars submissions

ABOUT SAHITHA RAO

10+ years of experience in end-to-end Data Engineering, Data Warehousing, and Big Data projects for enterprise clients.• Extensive hands-on experience with AWS services including S3, EMR, Glue, EC2, Redshift, and Kinesis for scalable cloud data architectures.• Proficient in designing, developing, and deploying ETL/ELT pipelines leveraging AWS Glue, Data Pipeline, and EMR.• Advanced skills in BigQuery, Dataproc, Cloud Data Fusion, and Cloud Storage for building modern data platforms on GCP.• Built and optimized machine learning solutions using Vertex AI, Dataproc, Python, and Spark for predictive analytics on Google Cloud.• Developed real-time streaming data integration and analytics using Spark Streaming, Pub/Sub, and Kafka on GCP-based architectures.• Designed and maintained data lakes and data warehouses on GCP utilizing dimensional modeling with Star and Snowflake schemas in BigQuery and Cloud Storage.• Expert in Spark, PySpark, and SQL for high-performance data processing, transformation, and ad hoc analytics.• Automated data workflows and pipeline orchestration with Apache Airflow, Oozie, and AWS-native tools.• Led successful migration projects from on-premises databases (Oracle, SQL Server, Teradata) to AWS and Azure platforms.• Authored and optimized complex SQL, PL/SQL, and Hive queries for business reporting and analytics.• Built CI/CD pipelines for automated deployment and testing of code and models using Azure DevOps, Jenkins, Docker, and Git.• Conducted advanced data profiling, cleansing, validation, and quality assurance for large and complex datasets.• Integrated and managed data from a wide range of sources including APIs, RDBMS, NoSQL, and cloud storage.• Delivered BI and reporting solutions leveraging Power BI and Snowflake for actionable business insights.• Implemented robust security, access control, and governance using AWS IAM, Key Vault, and Purview.• Used Terraform and CloudFormation for infrastructure as code, automating cloud resource provisioning and environment setup.• Collaborated cross-functionally with business, analytics, and data science teams to translate requirements into data products.• Integrated Databricks with Azure Synapse, Power BI, and Looker for seamless analytics and business intelligence workflows.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.