Jubayer Shaon

\"📊 Big Data Engineer with 5 Years Mastery | Sculpting Insights with Hadoop, Spark, AWS & Python 🚀\"

Role
Bigdata Engineer at Capital One
Location
Bronx, NY, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Jubayer Shaon

Big Data Engineer | 5+ Years of Deep Dive into Data Oceans Specialization: Design, implementation, and support of cutting-edge big data solutions leveraging Apache Spark, Hadoop, and AWS. Full-cycle project execution: from requirements gathering to architecture design, development, performance tuning, data extraction, cleansing, and insightful reporting. Tech Expertise: Big Data: Mastery over Spark SQL, Spark Streaming, Hadoop ecosystem (Hive, Hbase, Sqoop, Nifi, Kafka, Oozie, Cassandra, MapReduce, Flink), with hands-on experience in Cloudera & Hortonworks distributions. Cloud Platforms: Proficient with AWS offerings including EC2, S3, EMR, RDS, DynamoDB, Redshift, Athena, and more. Languages: Solid coding skills in Java, Scala, and Python. Databases: Command over Oracle, MySQL, PostgreSQL, SQL Server. Tools & IDEs: Familiarity with Airflow, Git, Scala IDE, IntelliJ, PyCharm. Industry Exposure: Made notable impacts across sectors including financial, banking, and telecom. Proud contributor to dynamic teams, including a significant stint at CapitalOne Bank. 🤝 Personal Traits: Recognized as a dedicated team player and an articulate communicator.

Experience

  1. Bigdata Engineer

    Capital One

    May 2020 — Present

    Collaborated with engineering teams to innovate solutions and foster knowledge sharing.Acted as a bridge with product teams, addressing vital customer-related challenges.Comprehensive experience in design, modeling, development, and testing.Authored and automated test scripts in Python, ensuring robust software delivery.Migrated on-prem images to the cloud via the Azure Copy toolkit.Mastered ETL processes with Airflow, creating pipelines and utilizing MapReduce for AWS S3 to BigQuery transfers.Managed tables in BigQuery and orchestrated data loads.Spearheaded the import/export of data between local systems, RDBMS, and HDFS.Demonstrated proficiency in deploying, configuring, and maintaining services on Azure Cloud, coupled with troubleshooting expertise.Handled large-scale data systems, including installation, upgrades, and migrations.Analyzed SQL scripts, transitioning solutions to Pyspark implementations.Developed performant Spark code with Python (Pyspark) and leveraged SparkAPI for Hive data analytics.Skilled in data cleaning and manipulation using advanced statistical tools.Comprehensive knowledge of Spark Architecture, Databricks, and Structured Streaming.Seamlessly integrated AWS and Azure with Databricks, managing clusters and overseeing the entire Machine Learning lifecycle.Automated workflows using Apache Airflow and shell scripting for consistent production execution.Regularly liaised with business units to align technical solutions with business needs.Actively participated in project planning, ensuring timely and quality delivery.

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.