Pancham Banerjee

Data Scientist @Pacific Data Integrators

Vancouver, BC, CA
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Feb 2020 — Present

Data Scientist @Pacific Data Integrators

View department →

Vancouver, BC, CA

Developed a multi-agent AI system (CrewAI + LangGraph) to monitor and auto-resolve or escalate interrupts in Informatica IDMC workflows, with guardrails for safe execution. Built supporting frontend and API routes in Next.js- Developed a custom code-conversion tool leveraging Retrieval-Augmented Generation (RAG) and Large Language models (Llama 3, Claude 4 Sonnet, GPT-4o)- Coded and deployed LLM-based customer service chatbots with RAG, integrating custom documents, vector databases (Pinecone, Chroma), and advanced prompt engineering. Delivered production-ready apps via Streamlit and Gradio, using both Llama 3 (open-source) and GPT-4o (closed-source) models- Designed and deployed a semantic product search engine in Databricks using Dolly (LLM), leveraging MLflow for end-to-end lifecycle management (tracking, versioning, and deployment), ensuring reproducibility and scalable delivery- Constructed industry-grade end-to-end Retail Recommender model using Databricks and PySpark, using both custom and AutoML frameworks for a dataset with users and products- Predicted outage causes for 1.2 million unclassified incidents (weather, equipment failure, tree-related etc) using supervised learning - Built a product-based Recommender Model for 1.5 million Electric Utility customers, determining the optimal energy utility product per customer (Autopay, e-bill, CARE, etc.)- Used LSTM networks to forecast Hourly Consumption and Charges for 2.7 million Residential Customers. Utilized classical forecasting methods as a baseline- Worked on Lifetime Value Modeling and Risk Score Analysis (Unsupervised Clustering) for 2.7 million Residential Customers, with the goal of estimating the Financial impact of COVID-19 on the customers and subsequently the company. Used an unsupervised clustering approach (Gaussian Mixture Models) and validated the clustering using supervised Gradient-Boosting Models (LightGBM).

EDUCATION

2018

Coursera

Deep Learning, a 5-course specialization by deeplearning.ai on Coursera

2007 — 2010

University of Calcutta

Bachelor of Science (B.Sc.), Physics

2012 — 2019

University of Southern California

Doctor of Philosophy - PhD, Astronomy and Astrophysics

2010 — 2012

Indian Institute of Technology, Kanpur

Master of Science (M.Sc.), Physics

2018

Coursera

Applied Data Science with Python, a 5-course specialization by University of Michigan on Coursera

2020

Coursera

TensorFlow in Practice, a 4-course specialization by deeplearning.ai on Coursera

SKILLS

ResearchMatlabData AnalysisStatisticsCC++LatexRPhysicsScienceFortranData VisualizationPythonSqlScikit-LearnMatplotlibGnuplotSeabornWeb ScrapingMachine LearningMicrosoft OfficeAstrophysicsComputational PhysicsProgramming

ABOUT PANCHAM BANERJEE

I currently work as a Data Scientist with Pacific Data Integrators. I completed my PhD from the University of Southern California in 2019, working primarily on Computational Astrophysics. Present obsessions include Large Language Models, Deep Reinforcement Learning, 2D Game Design and Electronic Music Production. Active projects- CitationVerify - Open Source CLI based tool for validating and downloading references in an Arxiv publicationProjects in the backlog- access-spotify: A Python package to query the spotify API for all the music data one could possibly want, to visualize, analyze etc etc -> https://github.com/panchambanerjee/access_spotify - convnet-feature-extractor: A Python package to extract and visualize features from different ConvNet architectures - fpl-data-analyzer: A Python package to download and analyze the data for the Fantasy Premier League- CosmologyAI: A suite of Data Science, Analytics and LLM-based projects on Cosmology and Extragalactic Astronomy- Ask-ArXiv, a research helper bot- Ad-LLama: For all your advertising needs!(Generate product names, slogans and logos.) I have experience in end-to-end software development, from data sourcing, to MongoDB storage, to building and testing predictive models, as well as Dockerization. I have also developed an extensive skill-set in conveying data science information from a business standpoint, via presentations and Tableau visualizations. I am a Kaggle Expert (both on Kernels and Discussions), currently (as of 2019) in the top 0.25% of the community.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Pancham Banerjee — Data Scientist at Pacific Data Integrators in Vancouver, BC, CA | Unifers