R Bunti

Spark/Hadoop Developer at Morgan Stanley

Role
Spark Hadoop Developer at Morgan Stanley
Location
New York, NY, US
LinkedIn followers
500 followers

About R Bunti

Environment: Java 1.8, Scala 2.10.5, Apache Spark 1.6.2, Apache Zeppelin, GreenPlum 4.3 (PostgreSQL), Treadmill, CDH 5.8.2, Spring 3.0.4, ivy 2.0, Gradle 2.13, Hive, HDFS, YARN, MapReduce, Sqoop 1.4.3, Flume, SOLR, HBase, Apache Cassandra, UNIX Shell Scripting, Python 2.6, AWS S3, Jenkins.

Experience

  1. Spark Hadoop Developer

    Morgan Stanley

    Oct 2016 — Present · New York, NY, US

    Responsibilities-•Creating end to end spark-Solr applications using Scala to perform various data cleansing, validation, transformation and summarization activities according to the requirement•Implemented moving average and regression analysis on input data.•Worked on Lily for indexing the data added/updated/deleted in HBase database to Solr collection. Indexing allows to query data stored in HBase with the Solr service. •Used Lilly Indexer for supporting flexible, custom, application-specific rules to extract, transform, and load HBase data into Solr.•Solr search results for columnFamily:qualifier links back to the data stored in HBase•Solr Schema field optimization with analyzer and filter based on requirement. •Data migration from various data sources to SOLR via stages according to the requirement•Responsible for loading customer\'s data and event logs into HBase using Scala API•Created HBase tables to store variable data formats of input data coming from different portfolios.•Implemented Moving averages, Interpolations and Regression analysis on input data.•Developed HQL queries for CRUD•Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data.•Performance tuning of long running Greenplum user defined functions. Leveraged the feature of temporary tables break the code into small sub part load to a temp table and join it later with the corresponding join tables. Table distribution keys are modified based on the data granularity and primary key column combination.•Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs. •Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.•Converted Ant application into Gradle

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

R Bunti — Spark Hadoop Developer at Morgan Stanley in New York, NY, US | Unifers