Pratyusha A
Data Engineer @CVS Health
Signup · Get unlimited contacts
WORK HISTORY
Data Engineer @CVS Health
Dallas, TX, US
Designed and developed AWS-native ETL pipelines using Glue, Lambda, and S3 for ingesting insurance datasets.• Migrated multiple on-prem pipelines into AWS Redshift and S3-based data lake, reducing processing latency.• Created Athena queries and Redshift Spectrum views for interactive analytics.• Implemented real-time ingestion using Kinesis Streams and AWS Lambda triggers.• Developed data lake partitions and Parquet formats in S3 for query performance optimization.• Applied KMS encryption and fine-grained IAM role policies for HIPAA-compliant PHI/PII data.• Automated infrastructure provisioning with CloudFormation and Terraform.• Designed CloudWatch dashboards and alarms for pipeline health monitoring.• Configured VPC networks and private endpoints to enhance cloud security.• Built serverless workflows with AWS Step Functions coordinating multiple Glue and Lambda jobs.• Optimized Redshift queries by applying distribution keys and sort keys.• Created CI/CD pipelines with Jenkins and GitHub for data engineering deployments.• Collaborated with data analysts to generate curated datasets for BI tools.• Implemented data quality checks in Python and Lambda to validate incoming files.• Partnered with business stakeholders to deliver datasets for claims processing, improving turnaround time.• Experience in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns. • Responsible for estimating the cluster size, monitoring, and troubleshooting of the Spark data bricks cluster. • Created Unix Shell scripts to automate the data load processes to the target Data Warehouse. • Responsible for implementing monitoring solutions in Ansible, Terraform, Docker, and Jenkins.
ABOUT PRATYUSHA A
8+ years professional experience in Big Data Development primarily using Hadoop and SparkEcosystems and design, development, and Implementation of Big data applications usingHadoop ecosystem frameworks and tools like HDFS, MapReduce, Yarn, Pig, Hive, Sqoop, Spark,Storm HBase, Kafka, Flume, Nifi, Impala, Oozie, Zookeeper, Airflow, etc.TECHNICAL SKILLS:Big Data Frameworks : Hadoop (HDFS, MapReduce), Spark, Spark SQL, SparkStreaming, Hive, Impala, Kafka, HBase, Flume, Pig, Sqoop,Oozie, Cassandra,AirflowBig Data Distribution : Cloudera, HortonworksProgramming Languages : Python, Java, Scala, Shell Scripting,PysparkOperating Systems : Windows, Linux (Ubuntu, Cent OS), Android, MacOSDatabases : Oracle, SQL Server, MySQL, Mongo DBCloud Technologies : AWS, GCP, AzureDesigning Tools : UML, VisioIDEs : Eclipse, NetBeansJava Technologies : JSP, JDBC, Servlets, JunitWeb Technologies : XML, HTML, JavaScript, jQuery, JSONLinux Experience: System Administration Tools, PuppetDevelopment methodologies: Agile, WaterfallLogging Tools: Log4jApplication / Web Servers : Apache Tomcat, WebSphereMessaging Services : ActiveMQ, Kafka, JMSVersion Tools: Git and CVSOthers : Putty, WinSCP, Data Lake, Talend, Terraform
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.