Samson Miah
Hadoop Developer (Data Engineer)
- Role
- Hadoop Developer (Data Engineer) at Bank of America
- Location
- Queens, NY, US
- LinkedIn followers
- 500 followers
About Samson Miah
Over 9+ years of IT experience as a Developer, Designer & quality Tester with cross platform integration experience using Hadoop Ecosystem. Strong knowledge of Hadoop Architecture and Daemons such as HDFS, JOB Tracker, Task Tracker, Name Node, Data Node and Map Reduce concepts. Well versed in implementing E2E solutions on big data using Hadoop framework. Worked with join patterns and implemented Map side joins and Reduce side joins using Map Reduce. Developed multiple MapReduce jobs to perform data cleaning and preprocessing. Having experience in developing a data pipeline using Kafka to store data into HDFS. Good knowledge on AWS infrastructure services Amazon S3, EMR, and Amazon Elastic Compute Cloud. Experience in composing shell scripts to dump the shared information from MySQL servers to HDFS. Worked on Implementing and optimizing Hadoop/MapReduce algorithms for Big Data analytics. Worked with Oozie and Zookeeper to manage the flow of jobs and coordination in the cluster. Experience in performance tuning, monitoring the Hadoop cluster by gathering and analyzing the existing infrastructure using Cloudera manager. Experience in building Azure stream Analytics ingestion spec for data ingestion which helps users to get sub second results in Real Time. Experience in building ETL(Azure Data Bricks) data pipelines leveraging Pyspark, Spark SQL. Experience in building the Orchestration on Azure Data Factory for scheduling purposes. Experience working with Azure Logic APP Integration tool. Expertise on working with databases like Azure SQL DB, Azure SQL DW. Hands - on experience in Azure Analytics Services - Azure Data Lake Store (ADLS), Azure Data Lake Analytics (ADLA), Azure SQL DW, Azure Data Factory (ADF), Azure Data Bricks (ADB) etc. Orchestrated data integration pipelines in ADF using various Activities like Get Metadata, Lookup, For Each, Wait, Execute Pipeline, Set Variable, Filter, until, etc. Programming experience working with Python, Scala. Extensively worked on Azure Databricks Experience in building the Orchestration on Azure Data Factory for scheduling purposes. Experience on Implementation of Azure log analytics providing Platform as a service for SD - WAN firewall logs. Experience in building the data pipeline by leveraging the Azure Data Factory. Selecting appropriate low cost driven AWS/Azure services to design and deploy an application based on given requirements. Solid programming experience on working with Python, Scala. Experience working in a cross-functional AGILE Scrum team.
Experience
Hadoop Developer (Data Engineer)
Jun 2021 — Present · Pennington, NJ, US
Worked on developing architecture document and proper guidelines. Worked with product owners, Designers, QA and other engineers in Agile development environment to deliver timely solutions to as per customer requirements. Hands on experience with Microsoft Azure Cloud services, Storage Accounts and Virtual Networks Worked on Microsoft Azure Storage - Storage accounts, blob storage, managed and unmanaged storages. Deployed Azure resource manager-based resources. Transform and analyze the data using Pyspark, Hive, based on ETL mapping. Developed Pyspark programs and created the data frames and worked on Transformation. Re-write some hive queries to spark SQL to reduce the overall batch Item. Completed data extraction, aggregation, and analysis in HDFS by using Pyspark and store the data needed to Hive. Involved in Migration projects to migrate data from data warehouses on Oracle/BD2 and migrated those to Teradata. Design and developed Batch processing and real-time processing solutions using ADF, Databricks clusters and stream Analytics. Created numerous pipelines in Azure using Azure Data Factory v2 to get the data from disparate source systems by using different Azure Activities like Move & Transform, Copy, filter, for each, Databricks etc. Maintain and provide support for optimal pipelines, data flows and complex data transformations and manipulations using ADF and PySpark with Databricks. Automated jobs using different triggers like Events, Schedules and Tumbling in ADF. Created, provisioned different Databricks clusters, notebooks, jobs and autoscaling. Performed data flow transformation using the data flow activity. Used Polybase to load tables in Azure synapse. Implemented Azure, self-hosted integration runtime in ADF. Improved performance by optimizing computing time to process the streaming data by optimizing the cluster run time. Scheduled, automated business processes and workflows using Azure Logic Apps.
Education
CUNY York College
Bachelor's degree in ( Accounting)
2001 — 2006
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.