Anuj Shahdeo

Data Engineer @ IBM | ML|Spark|AWS|Python|Java|SQL|Kafka|Airflow|Databricks|Scala|GCP|Hive|Hadoop|

Role
Lead Data Scientist at IBM
Location
Kolkata, WB, IN
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Anuj Shahdeo

Project: Release Risk Prediction SystemTech Stack: Python, R, Pandas, REST APIs, scikit-learn, XGBoost, Snowflake, FastAPI, Docker, Git, Jira, Gerrit, DToolCore Responsibilities: • Spearheaded the end-to-end migration of legacy R-based analytics pipelines to optimized, production-ready Python scripts, improving runtime efficiency and model maintainability. • Re-engineered 20+ complex R scripts involving statistical modeling, regression, and deployment KPIs into modular, testable Python modules using Pandas, NumPy, and scikit-learn. • Refactored legacy pin_write and dplyr transformations into Python equivalents using Polars, pyarrow, and Parquet, ensuring functional parity and performance gain. • Developed robust data extraction layers integrating Gerrit, Jira, DTool, and Confluence APIs using secure token-based authentication and batch-fetching logic. • Built unified, versioned data models combining code quality metrics, issue lifecycle patterns, and test/defect analytics stored in Snowflake and served via S3 for downstream ML consumption. • Trained predictive models to forecast risks such as: • Release delays • Deployment defects • Cost deviations • Patch failures during UAT • Tuned and validated models using grid search, AUC scoring, and business-aligned thresholds, boosting prediction accuracy and interpretability for non-technical users. • Exposed trained models via FastAPI microservices, containerized with Docker, and deployed in CI/CD pipelines using GitHub Actions and AWS Lambda/ECS. • Created monitoring tools to detect data drift, API failures, and model performance degradation, ensuring reliability across release cycles. • Partnered with QA and product teams to embed ML predictions into release dashboards, enhancing decision-making with real-time risk intelligence.With a Bachelor\'s in Computer Engineering from Rajiv Gandhi Proudyogiki Vishwavidyalaya, I bring structured thinking and a strong grasp of data structures to my role. My commitment to creating robust solutions is evident in the successful management of batch and real-time data processing requirements, ensuring our business stakeholders receive timely insights. In collaboration with cross-functional teams, we uphold high data quality and consistency, underpinning the organization\'s data-driven decision-making processes.

Experience

  1. Lead Data Scientist

    IBM

    Feb 2025 — Present · IN

    Develop and Deploy Machine Learning Models •Designed and implemented end-to-end machine learning pipelines for predictive analytics and classification problems. •Used supervised and unsupervised learning techniques, including regression, clustering, and deep learning. •Tuned hyperparameters and optimized models for accuracy, precision recall. •Big Data Processing and Analytics •Processed and analyzed large-scale datasets using PySpark and Apache Spark for efficient distributed computing. •Built scalable ETL pipelines to clean and transform raw data into structured formats. •Worked with Hadoop,Hive,SparkSQL to manage and query big data efficiently. •Cloud Computing with AWS •Developed and deployed machine learning models on AWS using SageMaker,Lambda, StepFunctions, and Glue. •Managed Amazon S3,DynamoDB and Redshift for data storage and retrieval. •Implemented CI/CD pipelines with AWS CodePipeline, CodeBuild, and CloudFormation. •Orchestrated large-scale data workflows using AWS Glue, EMR (Elastic MapReduce), and Kinesis for real-time and batch processing. •Data Engineering and Automation •Designed and implemented data pipelines for ingesting, processing, and storing high-volume structured and unstructured data. •Automated ETL workflows using Apache Airflow and AWS Glue workflows. •Optimised Spark jobs for high-performance data processing. •Statistical Analysis & Data Visualization •Performed exploratory data analysis (EDA) using Python (Pandas, NumPy, Matplotlib, Seaborn). •Created dashboards and reports using Tableau, Power BI, and QuickSight. •Implemented A/B testing and statistical methods to evaluate model effectiveness. •Performance Optimization and Model Monitoring •Deployed scalable models in production environments using Docker, Kubernetes, and AWS Fargate. •Monitored model drift and performance degradation, retraining models as needed. •Used MLflow and SageMaker Model Monitor for tracking experiments and model lifecycle management.

Education

  • Rajiv Gandhi Proudyogiki Vishwavidyalaya

    Bachelor's Degree, Computer Engineering

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Anuj Shahdeo — Lead Data Scientist at IBM in Kolkata, WB, IN | Unifers