Bhavneet Singh

Lead Ai Ml Engineer @National Research Council Canada / Conseil national de recherches Canada

Ottawa, ON, CA
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

May 2025 — Present

Lead Ai Ml Engineer @National Research Council Canada / Conseil national de recherches Canada

View department →

Ottawa, ON, CA

Architected a Graph-RAG agentic system using frontier LLMs and LangChain, with distributed inference over 150K+ documents.• Built multi-agent orchestration (LangGraph + CrewAI + MCP + Ray) with policy routing, improving decision accuracy by 26%.• Implemented hybrid retrieval (MiniLM-L6-v2 + ColBERT) with feedback loops, improving retrieval precision@k by 31%.• Optimized FAISS, Milvus, and PostgreSQL vector stores to index 3M+ legal docs with sub-second retrieval latency.• Deployed LLM inference on vLLM + Triton, enabling multi-tenant serving with predictable latency/memory-efficient utilization.• Designed LLM microservices (FastAPI + ONNX + CUDA graphs) achieving <400ms inference under concurrent production load.• Applied VLMs (CLIP, BLIP-2) with hybrid retrieval and contextual reasoning, improving real-time inference accuracy by 18%.• Fine-tuned LLMs/vision models using LoRA-based tuning and quantization, boosting accuracy by 24% and reducing size by 48%.• Developed distributed ML pipelines (PySpark + Spark + PyTorch) over multi-node clusters, improving throughput by 65%.• Built real-time ingestion and embedding pipelines with PII filtering/schema validation, accelerating ETL throughput by 52%.• Integrated LangSmith + MLflow + W&B observability, reducing debugging cycles by 40% and standardizing RAG quality metrics.• Automated drift detection, retraining loops, and lineage tracking via Databricks, reducing performance degradation by 33%.• Implemented token tracking, prompt versioning, and cost monitoring to control LLM spend and predictable production scaling.

EDUCATION

N/A

University of Ottawa

Master of Computer Science

ABOUT BHAVNEET SINGH

I am an AI Engineer specializing in building and operating production-scale Generative AI and machine learning systems, with expertise in LLMs, RAG, and distributed inference. My focus is on transforming advanced AI models into scalable, low-latency systems that deliver measurable real-world impact.At the National Research Council Canada (NRC), I architect and deploy production AI platforms:• Built a Graph-RAG regulatory intelligence system enabling semantic search across 150K+ documents with sub-400 ms inference latency, improving retrieval precision by 31%• Developed distributed LLM microservices using FastAPI, vector databases, and cloud infrastructure (AWS, GCP), enabling scalable real-time inference• Designed multimodal AI pipelines achieving sub-300 ms real-time inference, optimized for production reliability and performanceI specialize in designing end-to-end AI systems integrating model development, deployment, and infrastructure using PyTorch, HuggingFace, Docker, Kubernetes, MLflow, and cloud platforms (AWS, GCP, Azure), with a strong focus on scalability, observability, and production reliability.I am also the creator of RiskReg, an open-source rare-event regression framework that reduced prediction error by up to 3× across 34 real-world datasets.I focus on building production-ready AI systems that combine LLMs, MLOps, and scalable infrastructure to deliver reliable, high-performance real-world solutions.

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Bhavneet Singh — Lead Ai Ml Engineer at National Research Council Canada / Conseil national de recherches Canada in Ottawa, ON, CA | Unifers