Tushib Haque
IT Professional | 11+ Years Experience | Big Data, Data engineering, AWS & Azure Cloud | Multi-Tenant Data Environments
- Role
- Data Engineer at Exelon
- Location
- Allentown, PA, US
- LinkedIn followers
- 500 followers
Experience
Data Engineer
Nov 2021 — Present
Responsibilities & Accomplishments:We are building Hadoop cluster to support the existing business systems. Development process is building solutions for batch data, as well as real time data ingestions. As ecosystem, we are using HIVE, HBase, and Impala.· Installed, configured and maintained Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Oozie, Flume, Zookeeper and Sqoop.· Installing and Upgrading Cloudera CDP 7.1.9 on production.· Moving the Services (Re-distribution) from one Host to another host within the Cluster to facilitate securing the cluster and ensuring High availability of the services.· Installed and configured Hadoop, MapReduce, HDFS (Hadoop Distributed File System), developed multiple MapReduce jobs in java for data cleaning.· Worked on installing cluster, commissioning & decommissioning of Data Nodes, Name Node recovery and capacity planning.· Cloudera Data Lake environments for high availability and scalability setup· Enabled Oozie to schedule workflows relating to hive and other jobs.· Used Sqoop to import and export data from HDFS to RDBMS and vice-versa.· Created Hive tables and involved in data loading and writing Hive UDFs.· Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.· Worked on HBase, Hive, and Impala.· Automated workflows using shell scripts to pull data from various databases into Hadoop.· Deployed Hadoop Cluster in Fully Distributed and Pseudo-distributed modes.· Used Nagios and Ganglia for monitoring tools· Designed and implemented scalable ETL pipelines in Azure Data Factory to automate data extraction, transformation, and loading for real-time reporting in healthcare environments.· PySpark and SQL to process large datasets, improving data processing.· Azure Synapse and Databricks for seamless data warehousing and analytics, enabling faster decision-making across multiple departments.· custom data models in Azure Synapse, optim
Skills
- Sql
- Microsoft Office
- Leadership
- Microsoft Word
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.