Vineeth Sanke

AI/ML Engineer | LLMs, RAG, Fine-tuning (LoRA, QLoRA) | Agents (CrewAI) | LangChain, Vector DBs | Bedrock, NVIDIA NIM, AWS | NLP & Deep Learning (Transformers)

Role
Ai Ml Engineer at Spring Health
Location
Jersey City, NJ, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Vineeth Sanke

I’m a AI/ML Engineer and Machine Learning Engineer specializing in Generative AI, NLP, and scalable data systems, with 5+ years of experience building production-grade ML and GenAI solutions across healthcare and financial services. My work focuses on designing and deploying end-to-end AI systems from data ingestion and feature engineering to model training, fine-tuning, evaluation, and large-scale deployment. I have hands-on experience building LLM-powered applications, including Retrieval-Augmented Generation (RAG) pipelines, transformer-based architectures, and multi-agent AI systems for real-world use cases. At Spring Health, I developed scalable ML and GenAI pipelines processing millions of records to support clinical and financial reporting. I built robust RAG-based architectures with strong guardrails to improve response accuracy, reduce hallucinations, and ensure reliability in sensitive healthcare environments. Core Expertise: • Generative AI & LLMs (OpenAI, Hugging Face, Ollama, Amazon Bedrock, NVIDIA NIM) • Fine-tuning LLMs (LoRA, QLoRA) • RAG Pipelines, LangChain, Embeddings, Hybrid Search • Multi-Agent Systems (CrewAI, LangGraph) • NLP & Deep Learning (Transformers, LSTM, GRU, ANN) • Machine Learning & Predictive Modeling • Data Engineering (PySpark, Airflow, Snowflake) • Cloud Platforms (AWS, Azure) I’m passionate about building AI systems that are not only powerful, but also scalable, reliable, and production-ready.

Experience

  1. Ai Ml Engineer

    Spring Health

    Sep 2024 — Present · NY, US

    Scalable LLM and RAG Infrastructure: Designed and implemented high-throughput data ingestion pipelines using AWS Glue and Amazon SageMaker, processing 5M+ records into optimized Parquet formats to power Retrieval-Augmented Generation (RAG) systems, reducing LLM inference latency by 50%. • LLM Data Pipeline and Embedding Engineering: Built end-to-end NLP preprocessing and embedding pipelines, transforming large-scale unstructured clinical text into tokenized, vectorized, and embedding-ready datasets, reducing data preparation time by 40% and enabling efficient semantic retrieval. • Production-Grade AI API Development: Developed scalable RESTful APIs using FastAPI to serve LLM and SageMaker endpoints, implementing semantic caching and optimized retrieval strategies to improve response time by 35% and reduce compute costs. • Generative AI MLOps and Orchestration: Architected automated CI/CD pipelines for LLM workflows using SageMaker Pipelines and AWS Step Functions, ensuring reproducibility of prompts, embeddings, datasets, and model outputs across environments. • Vector Database and Feature Store Integration: Designed and maintained a centralized feature store and embedding repository (SageMaker Feature Store), enabling efficient storage, retrieval, and reuse of vector embeddings for RAG-based applications. • AI Observability, Guardrails and Monitoring: Implemented LLM monitoring and governance using Python and SageMaker to detect data drift, schema anomalies, bias, and hallucinations, integrating webhook-based alerting systems for real-time model reliability in production environments.

Education

  • Pace University

    Master's degree, Computer Science

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Vineeth Sanke — Ai Ml Engineer at Spring Health in Jersey City, NJ, US | Unifers