Jubayer Shaon
\"📊 Big Data Engineer with 5 Years Mastery | Sculpting Insights with Hadoop, Spark, AWS & Python 🚀\"
- Role
- Bigdata Engineer at Capital One
- Location
- Bronx, NY, US
- LinkedIn followers
- 500 followers
About Jubayer Shaon
Big Data Engineer | 5+ Years of Deep Dive into Data Oceans Specialization: Design, implementation, and support of cutting-edge big data solutions leveraging Apache Spark, Hadoop, and AWS. Full-cycle project execution: from requirements gathering to architecture design, development, performance tuning, data extraction, cleansing, and insightful reporting. Tech Expertise: Big Data: Mastery over Spark SQL, Spark Streaming, Hadoop ecosystem (Hive, Hbase, Sqoop, Nifi, Kafka, Oozie, Cassandra, MapReduce, Flink), with hands-on experience in Cloudera & Hortonworks distributions. Cloud Platforms: Proficient with AWS offerings including EC2, S3, EMR, RDS, DynamoDB, Redshift, Athena, and more. Languages: Solid coding skills in Java, Scala, and Python. Databases: Command over Oracle, MySQL, PostgreSQL, SQL Server. Tools & IDEs: Familiarity with Airflow, Git, Scala IDE, IntelliJ, PyCharm. Industry Exposure: Made notable impacts across sectors including financial, banking, and telecom. Proud contributor to dynamic teams, including a significant stint at CapitalOne Bank. 🤝 Personal Traits: Recognized as a dedicated team player and an articulate communicator.
Experience
Bigdata Engineer
May 2020 — Present
Collaborated with engineering teams to innovate solutions and foster knowledge sharing.Acted as a bridge with product teams, addressing vital customer-related challenges.Comprehensive experience in design, modeling, development, and testing.Authored and automated test scripts in Python, ensuring robust software delivery.Migrated on-prem images to the cloud via the Azure Copy toolkit.Mastered ETL processes with Airflow, creating pipelines and utilizing MapReduce for AWS S3 to BigQuery transfers.Managed tables in BigQuery and orchestrated data loads.Spearheaded the import/export of data between local systems, RDBMS, and HDFS.Demonstrated proficiency in deploying, configuring, and maintaining services on Azure Cloud, coupled with troubleshooting expertise.Handled large-scale data systems, including installation, upgrades, and migrations.Analyzed SQL scripts, transitioning solutions to Pyspark implementations.Developed performant Spark code with Python (Pyspark) and leveraged SparkAPI for Hive data analytics.Skilled in data cleaning and manipulation using advanced statistical tools.Comprehensive knowledge of Spark Architecture, Databricks, and Structured Streaming.Seamlessly integrated AWS and Azure with Databricks, managing clusters and overseeing the entire Machine Learning lifecycle.Automated workflows using Apache Airflow and shell scripting for consistent production execution.Regularly liaised with business units to align technical solutions with business needs.Actively participated in project planning, ensuring timely and quality delivery.
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.