Rithick Bisher
Senior Data Engineer at DuPont
- Role
- Senior Data Engineer at DuPont
- Location
- Memphis, TN, US
- LinkedIn followers
- 500 followers
About Rithick Bisher
6 years of IT experience in a variety of industries working on Big Data technology using technologies such as Cloudera and Hortonworks distributions. Hadoop working environment includes Hadoop, Spark, MapReduce, Kafka, Hive, Ambari, Sqoop, HBase, and Impala.• Fluent programming experience with Scala, Java, Python, SQL, T - SQL, R. • Hands-on experience in developing and deploying enterprise-based applications using major Hadoop ecosystem components like MapReduce, YARN, Hive, HBase, Flume, Sqoop, Spark MLlib, Spark GraphX, Spark SQL, Kafka.• Adept at configuring and installing Hadoop/Spark Ecosystem Components.• Proficient with Spark Core, Spark SQL, Spark MLlib, Spark GraphX and Spark Streaming for processing and transforming complex data using in-memory computing capabilities written in Scala. Worked with Spark to improve efficiency of existing algorithms using Spark Context, Spark SQL, Spark MLlib, Data Frame, Pair RDD’s and Spark YARN.• Experience in application of various data sources like Oracle SE2, SQL Server, Flat Files and Unstructured files into a data warehouse.• Able to use Sqoop to migrate data between RDBMS, NoSQL databases and HDFS.• Experience in Extraction, Transformation and Loading (ETL) data from various sources into Data Warehouses, as well as data processing like collecting, aggregating and moving data from various sources using Apache Flume, Kafka, PowerBI and Microsoft SSIS.• Hands-on experience with Hadoop architecture and various components such as Hadoop File System HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Hadoop MapReduce programming.• Comprehensive experience in developing simple to complex Map reduce and Streaming jobs using Scala and Java for data cleansing, filtering and data aggregation. Also possess detailed knowledge of MapReduce framework.• Used IDEs like Eclipse, IntelliJ IDE, PyCharm IDE, Notepad ++, and Visual Studio for development.• Seasoned practice in Machine Learning algorithms and Predictive Modeling such as Linear Regression, Logistic Regression, Bayes, Decision Tree, Random Forest, KNN, Neural Networks, and K-means Clustering.• Ample knowledge of data architecture including data ingestion pipeline design, Hadoop/Spark architecture, data modeling, data mining, machine learning and advanced data processing.• Experience working with NoSQL databases like Cassandra and HBase and developed real-time read/write access to very large datasets via HBase.
Experience
Senior Data Engineer
Jan 2024 — Present
Extensively involved in installation and configuration of Cloudera Distribution Hadoop platform.· Extract, transform, and load (ETL) data from multiple federated data sources (JSON, relational database, etc.) with Data Frames in Spark.· Utilized SparkSQL to extract and process data by parsing using Datasets or RDDs in Hive Context, with transformations and actions (map, flat Map, filter, reduce, reduce By Key).· Extended the capabilities of Data Frames using User Defined Functions in and Scala.· Resolved missing fields in Data Frame rows using filtering and imputation.· Integrated visualizations into a Spark application using Databricks and popular visualization libraries (ggplot, MatPlotLib).· Trained analytical models with Spark ML estimators including linear regression, decision trees, logistic regression, and k-means.· Performed pre-processing on a dataset prior to training, including standardization, normalization.· Created pipelines to create a processing pipeline including transformations, estimations, evaluation of analytical models.· Evaluated model accuracy by dividing data into training and test datasets and computing metrics using evaluators.· Tuned training hyper-parameters by integrating cross-validation into pipelines.· Computed using Spark MLlib functionality that wasn’t present in SparkML by converting DataFrames to RDDs and applying RDD transformations and actions.· Troubleshot and tuned machine learning algorithms in Spark.
Education
Trine University
Masters in Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.