Kaushik Nishtala
Lead Data Scientist | NLP, LLM Optimization & Applied ML Systems
- Role
- Lead Data Scientist at HiLabs
- Location
- Washington, DC, US
- LinkedIn followers
- 500 followers
About Kaushik Nishtala
Lead Data Scientist with > 5 YOE building scalable AI solutions across healthcare, IT, and early-stage startups - turning messy, high-stakes data problems into production systems that ship, scale, and stick.I build end-to-end AI platforms spanning policy optimized agentic systems, layout-aware document understanding, representation learning & embedding models, anomaly detection, and LLM fine-tuning & alignment (SFT/RLHF-style optimization), with emphasis on offline/online evaluation, reliability, model risk controls, and MLOps across AWS/Azure/GCP.To deep dive into my work, side projects, and technical notes: https://sosayskaushik.dev
Experience
Lead Data Scientist
Mar 2022 — Present · Washington, DC, US
Contracts Intelligence - NegotiateAI: Architected an LLM/VLM-powered negotiation assistant that extracts and compares contract clauses and rate tables against playbooks/templates with source-grounded deltas, achieving 85–92% precision/recall on change detection- DocuDelta: Productionized a post-hoc contract comparison service that performs structure reconstruction, clause alignment, and semantic diffing to generate audit-ready change logs, reducing review effort from ∼20 hours to ∼3 hours per contract- DocumentAI: Trained and deployed a custom Swin Transformer (DONUT encoder)+ Faster R-CNN (FPN) layout model for bounding-box chunking, achieving 96% mAP, and improved table/image-heavy retrieval with ColPali-style multimodal search- Clinical Analytics Chat Assist (KG-RAG + Text-to-SQL)- Delivered a GenAI clinical Q&A system that answers encounter-level and longitudinal questions using knowledge-graph hierarchical RAG and a self-correcting Text-to-SQL pipeline with constrained decoding for safe, reliable queries on regulated clinical data- Directed ML anomaly detection design, enhancing healthcare provider directory accuracy and saving $10M in claims processing- Improved data pipeline efficiency using Hadoop and Spark, and migrated infrastructure to AWS to bolster downstream operations- Established robust, secure infrastructure using self-hosted LLMs and asynchronous SQS process- Developed a scalable system for processing provider contracts, focusing on entity extraction to enhance claims processing- Engineered a document processing pipeline with LayoutLM for extracting entities from diverse media types- Served as primary technical liaison to clients like Elevance Health, translating complex algorithms into actionable business insights.
Education
Northeastern University
Master of Science, Data Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.