R Bunti
Spark/Hadoop Developer at Morgan Stanley
- Role
- Spark Hadoop Developer at Morgan Stanley
- Location
- New York, NY, US
- LinkedIn followers
- 500 followers
About R Bunti
Environment: Java 1.8, Scala 2.10.5, Apache Spark 1.6.2, Apache Zeppelin, GreenPlum 4.3 (PostgreSQL), Treadmill, CDH 5.8.2, Spring 3.0.4, ivy 2.0, Gradle 2.13, Hive, HDFS, YARN, MapReduce, Sqoop 1.4.3, Flume, SOLR, HBase, Apache Cassandra, UNIX Shell Scripting, Python 2.6, AWS S3, Jenkins.
Experience
Spark Hadoop Developer
Oct 2016 — Present · New York, NY, US
Responsibilities-•Creating end to end spark-Solr applications using Scala to perform various data cleansing, validation, transformation and summarization activities according to the requirement•Implemented moving average and regression analysis on input data.•Worked on Lily for indexing the data added/updated/deleted in HBase database to Solr collection. Indexing allows to query data stored in HBase with the Solr service. •Used Lilly Indexer for supporting flexible, custom, application-specific rules to extract, transform, and load HBase data into Solr.•Solr search results for columnFamily:qualifier links back to the data stored in HBase•Solr Schema field optimization with analyzer and filter based on requirement. •Data migration from various data sources to SOLR via stages according to the requirement•Responsible for loading customer\'s data and event logs into HBase using Scala API•Created HBase tables to store variable data formats of input data coming from different portfolios.•Implemented Moving averages, Interpolations and Regression analysis on input data.•Developed HQL queries for CRUD•Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data.•Performance tuning of long running Greenplum user defined functions. Leveraged the feature of temporary tables break the code into small sub part load to a temp table and join it later with the corresponding join tables. Table distribution keys are modified based on the data granularity and primary key column combination.•Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs. •Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.•Converted Ant application into Gradle
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.