Balaji Koneti

Senior Generative Ai Developer @Nordstrom

Denton, TX, US
MOBILE NUMBERS
+91 *********19

Signup · Get unlimited contacts

WORK HISTORY

Aug 2025 — Present

Senior Generative Ai Developer @Nordstrom

View department →

Plano, TX, US

Architected RAG-based LLM services using LangChain + pgvector, improving retrieval relevance by 22%(Precision@5) across ~450 real enterprise queries through semantic chunking and hybrid retrieval.• Reduced P95 end-to-end response latency from 1.3s to 640ms by introducing response caching, request batching, and separating embedding pipelines from generation services.• Decreased average LLM cost per request by 31% via token budgeting, prompt compression, and dynamic routing of retrieval-only queries to smaller models.• Built a scalable evaluation pipeline combining human-in-the-loop review (Label Studio) with LLM-as-a-judge grading (custom GPT-based evaluators) to detect faithfulness and relevance regressions after index or prompt updates.• Designed a scalable inference microservice using FastAPI and AWS, incorporating health checks, circuit-breaker patterns, and explicit handling for empty or low-confidence retrieval to ensure resilience under load.

EDUCATION

N/A

Northern Arizona University

Master of Science - MS, Computer Science

N/A

Jawaharlal Nehru Technological University, Anantapur

Bachelor of Technology - BTech, Computer Science

ABOUT BALAJI KONETI

I spent 4 years making ML reliable in production. Now I\'m doing the same for Gen AI and the results are measurable. At Nordstorm, I architect RAG systems and LLM inference pipelines for enterprise use. In 7 months of production work: · P95 latency: 1.3s → 640ms (51% reduction) · Cost per LLM request: down 31% via token budgeting & smart routing · Retrieval precision:+22% through semantic chunking & hybrid search · Evaluation: built LLM-as-a-judge pipelines that catch regressions before they shipBefore Gen AI, I built the ML foundations that make these systems possibleAt Infosys (2022–2023): ML fraud detection (~25% fraudulent activity), NLP claims automation (+35% productivity), ETL pipelines processing 1M+ daily records, TensorFlow models at 98% uptime all in financial & healthcare systems where failure isn\'t an option. At Nerds & Geeks (2020–2022): Real-time recommendation APIs, BI dashboards for 3 business units, MongoDB optimization (−40% query latency). I also taught what I know mentored 20+ graduate students in LLMs at Northern Arizona University (3.91 GPA, MS Computer Science), designed transformer fine-tuning curricula, and guided teams to 30% faster training iteration. The through-line across all of it: I care obsessively about ML systems that work in the real world not in notebooks. Stack: LangChain · pgvector · FastAPI · RAG · LLMOps · LLM-as-a-judge · PyTorch · TensorFlow · AWS (Certified ML Specialist) · Python · Airflow · Docker

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.