Moiez Mohammed
Hadoop/Spark Developer at Iqvia
- Role
- Data Engineer at IQVIA
- Location
- Tampa, FL, US
- LinkedIn followers
- 500 followers
About Moiez Mohammed
Data Engineer with 7 plus years of experience in interpreting and analyzing sophisticated datasets and expertise in providing business insights using the HADOOP ecosystem.• Strong working experience with ingestion, storage, processing, and analysis of Big Data. • Successfully loaded files to HDFS from Oracle, Sql Server and Teradata using Sqoop.• Working with Sqoop in importing and exporting data from different databases like MySQL, Oracle into HDFS and Hive. • Experience with working of cloud configuration in (Amazon web services) AWS. • Demonstrated experience in delivering data and analytic solutions leveraging AWS, Azure, or similar cloud data lake.• Experience on working structured, unstructured data with various file formats such as Avro data files, xml files, JSON files, sequence files, ORC and Parquet. • Experience with Oozie Workflow Engine to automate and parallelize Hadoop, MapReduce, and Pig jobs.• Expertise in pyspark and proficiency in working with distributed computing frameworks.• Expertise in developing Scala and Python applications and good working knowledge of working with Java.• Experience in working with databases, such as Oracle, SQL Server, My SQL. • Extensive experience with ETL and Query tools for Big Data like Pig Latin and HiveQL. • Extensive experience in ETL process consisting of data transformation, data sourcing, mapping, conversion, and loading.• Experience in developing data pipeline using Kafka, Spark, and Hive to ingest, transform and analyzing data. • Experience in Data modeling and connecting Cassandra from Spark and saving summarized data frame to Cassandra. • Exploring with the Spark for improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD\'s, Spark Yarn. • Developing applications using Scala, Spark SQL and MLLib libraries along with Kafka and other tools as per requirement then deployed on the Yarn cluster. • Adequate knowledge and working experience in Agile & Waterfall methodologies. • Developing and Maintaining the Web Applications using the Web server Tomcat, IBM WebSphere.• Experience in job workflow scheduling and monitoring tools like Oozie, Nifi.• Experience in Front-end Technologies like Html, CSS, Html5, CSS3, and Ajax. • Experience in building high performance and scalable solutions using various Hadoop ecosystem tools like Pig, Hive, Sqoop, Spark, Solr and Kafka.
Experience
Data Engineer
Jan 2022 — Present · NC, US
Developed, tested, and deployed Spark applications, ensuring efficient and reliable data processing.• Troubleshot and fine-tuned Spark jobs to optimize performance and address any issues or bottlenecks.• Built real-time pipelines using Kafka and Spark Streaming, enabling the processing of data streams in a timely manner.• Utilized Spark JDBC Readers to connect to external databases and extract data for storage in the S3 data lake.• Leveraged Spark JDBC Writers to connect to Redshift and write processed data frames directly into the Redshift database.• Employed Hive scripting to generate custom and adhoc datasets requested by downstream business teams, facilitating their data analysis needs.• Utilized the Glue metastore service in AWS to store and manage all Hive metadata, ensuring consistency and accessibility.• Utilized the Athena Interactive Query Service in AWS to perform advanced data analysis and exploration on stored datasets.• Automated the launching and termination of EMR Spark clusters using the AWS Java SDK, improving the scalability and cost-effectiveness of the infrastructure.• Developed Spring Boot-based REST applications to provide downstream application teams with access to metadata and previews of processed data.• Automated CICD (Continuous Integration and Continuous Deployment) build and deployment processes using Jenkins, streamlining the development workflow.• Built different modules using Scala Spark to perform various transformations on Spark DataFrames and store the processed data in the Parquet format.• Conducted a proof-of-concept (POC) for developing pipelines in Azure Synapse and Azure Data Factory, enabling the ingestion of data from external databases and loading it into Azure Blob containers. Utilized the Snowflake Snowpark API and developed a Scala-based project capable of ingesting various file types, connecting to different RDBMS systems, and loading the data into both Snowflake and other SFTP destinations.
Education
Osmania University
Bachelor of Technology - BTech
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.