Priyanka Bajaj
Distinguished Engineer - Ai Engineering @UBS
Signup · Get unlimited contacts
WORK HISTORY
Distinguished Engineer - Ai Engineering @UBS
Architected UBS\'s firm-wide LLM execution platform — the default ML layer for product teams, engineers, and workflows serving tens of millions of clients globally across regulated banking. • Reduced P99 inference latency by 32% through multi-model routing, dynamic batching, and caching — enabling LLM adoption into latency-sensitive, customer-facing workflows for the first time • Cut duplicated LLM development effort by 60–70% by centralising ML primitives (evaluation, retrieval, routing, safety, telemetry) while preserving model heterogeneity at the edges • Introduced explicit per-use-case failure budgets — shifting 120+ engineers from model-accuracy thinking to system-level reliability reasoning under uncertainty • Reduced compliance escalations measurably by embedding safety, auditability, and regulatory controls directly into the platform layer — enabling LLM deployment into regulated, client-facing paths • Mentored 10+ Staff and Principal engineers to independently validate ML architecture decisions; published LLM failure taxonomy postmortems adopted as firm-wide internal standards
EDUCATION
Liverpool John Moores University
Masters Degree - MSc in AI & ML
Devi Ahilya Vishwavidyalaya
Master of Business Administration - MBA
Banasthali Vidyapith
Bachelor of Science - BSc
ABOUT PRIYANKA BAJAJ
Most teams treat LLMs like deterministic services. At scale, that assumption becomes a liability. At UBS, I led the design and execution of a firm-wide LLM platform — the default ML execution layer for product teams, engineers, and workflows reaching tens of millions of clients globally. One of the largest regulated deployments of generative AI in global banking. The hard problem wasn\'t the models. It was redefining about correctness, safety, and cost in probabilistic systems. I centralised ML primitives — evaluation, retrieval, routing, safety, telemetry — while preserving model heterogeneity at the edges. Built on PyTorch, vLLM, and Kubernetes, with custom multi-model routing and evaluation pipelines measuring distributional drift, not just point accuracy. I introduced explicit failure budgets per use case, not per model. That shift moved teams from debating \"best model\" to reasoning about acceptable risk envelopes. Outcomes: → 32% P99 latency improvement across the platform → 60–70% reduction in duplicated LLM development effort → Onboarding time reduced from months to weeks → Measurable reduction in compliance escalations across regulated workflows → Unlocked LLM adoption into customer-facing and decision-critical paths Beyond the platform: I mentored 10+ Staff and Principal engineers to independently validate ML architecture decisions — and published postmortems on LLM failure taxonomy that became internal standards across the firm. I operate at the intersection of ML architecture, organisational design, and long-term system reliability. My decisions carry multi-year blast radius — and I design for that from day one.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.