Maina K
Data Engineer @CIBC
Signup · Get unlimited contacts
WORK HISTORY
Data Engineer @CIBC
Designed and implemented scalable big data pipelines using Hadoop, HDFS, Spark, Hive, Sqoop, and Oozie to process structured and unstructured customer data. Developed PySpark and Scala solutions to optimize and migrate legacy MapReduce jobs, improving data processing performance and scalability. Built a Snowflake data ingestion framework for both batch and real-time data using Snowflake Stages and Data Pipes, enabling seamless cloud migration from SQL Server and Oracle via Azure DMS. Engineered real-time streaming pipelines with Spark Streaming and Kafka, processing event logs and loading data into Cassandra for analytics. Automated ETL workflows with shell scripting and Python, integrating data from multiple databases and servers into Hadoop, improving efficiency and reducing manual intervention. Troubleshot and optimized Hive, Pig, and Spark jobs, resolving performance bottlenecks in joins, aggregations, and transformations. Configured and maintained Cloudera Hadoop ecosystem components (Flume, Hive, Pig, Sqoop, Oozie) and created workflows for scheduling and orchestration of ETL processes.
EDUCATION
Osmania University
Integrated MBA, Integrated MBA
ABOUT MAINA K
I am a Data Engineer with years of experience building and optimizing scalable data…
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.