Joseph Banna

Hadoop Developer at General Motors

Role
Hadoop Developer at General Motors
Location
Bronx, NY, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Joseph Banna

Experience in Big Data and Hadoop Administration ecosystems including HDFS, Pig, Hive, Impala, HBase, Yarn, Sqoop, Flume, Oozie, Hue, MapReduce, and Spark Worked with Play framework and Akka parallel processing. Expertise in deployment of Hadoop, Yarn, Spark integration with Cassandra, etc.Contributed to partitioning and bucketing of data, designed, and managed data and created external tables in Hive to optimize performance Experience in importing and exporting data between HDFS and RDBMS using Sqoop Strong experience and knowledge of real time data analytics using Spark, Kafka, and Flume Developing and maintaining Workflow Scheduling Jobs in Oozie for importing data from RDBMS to Hive. Utilized Spark Core, Spark Streaming and Spark SQL API for faster processing of data instead of using MapReduce in Java. Responsible for data extraction and data integration from different data sources into Hadoop Data Lake by creating ETL pipelines Using Spark, MapReduce and Hive. Involved in converting Hive/SQL queries into Spark transformations using Spark Data frames and Scala.Involved in streaming data from DB2 to HBase table using Spark Streaming and Apache Kafka. Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data. Developed Spark programs with Scala.Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Scala, and Python. Analyzed the SQL scripts and designed the solution to implement using Spark. Involved in converting MapReduce programs into Spark transformations using Spark RDD in Scala.Experienced in Apache Spark for implementing advanced procedures like text analytics and processing using the in-memory computing capabilities written in Scala. Hands on with real time data processing using distributed technologies Storm and Kafka. Experienced in writing queries and sub-queries for SQL, Hive, Impala and Spark; and used different Spark modules like Spark Core, Spark RDDs, Spark Data frame and Spark SQL. Experienced in converting Hive queries into Spark Transformations and Actions Worked on data serialization formats for converting complex objects into sequence bits by using CSS, Avro, Parquet, JSON and CSV. Strong command over relational databases including MySQL, Oracle, MS SQL Server, and MS Access Experience in providing good production support for 24x7 over weekends on rotation basis.

Experience

  1. Hadoop Developer

    General Motors

    Apr 2017 — Present · Detroit, MI, US

    Experienced in developing Spark scripts for data analysis in Scala.Used Spark-Streaming APIs to perform necessary transformations.Involved in converting Hive/SQL queries into Spark transformations using Spark SQL and Scala.Worked with spark to consume data from Kafka and convert that to common format using Scala.Converted existing MapReduce jobs into Spark transformations and actions using Spark RDDs, Data frames and Spark SQL APIs.Wrote new spark jobs in Scala to analyze the data of the customers and sales history.Involved in requirement analysis, design, coding, and implementation phases of the project.Used Spark API over Hadoop YARN to perform analytics on data in Hive.Experience in both SQL Context and Spark Session.Developed Scala based Spark applications for performing data cleansing and data aggregation,Worked on troubleshooting spark application to make them more error tolerant.Involved in HDFS maintenance and loading of structured and unstructured data and imported data from dataset to HDFS using Sqoop and written the Spark Script to process the HDFS data.Used Spark API over Hadoop YARN to perform analytics on data in Hive.Involved in Spark and Spark Streaming creating RDD\'s, applying operations -Transformation and Actions.Created partitioned tables and loaded data using both static partition and dynamic partition method.Implemented POC\'s on migrating to Spark-Streaming to process the live data.Executed Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business requirements.Ingested data from RDBMS and performed data transformations, and then export the transformed data to HDFS as per the business requirement.Used Impala to read, write and query the data in HDFS.Stored the output files for export onto HDFS and later these files are picked up by downstream systems.Load the data into Spark RDD and do in memory data Computation to generate the Output response.

Education

  • BL College Khulna

    Bachelor of Arts - BA, English Language and Literature, General

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Joseph Banna — Hadoop Developer at General Motors in Bronx, NY, US | Unifers