Harsh Kumar

Turning LLMs into Production-Grade AI | AI/ML Engineer | RAG · Agentic Systems · Fine-Tuning · Full-Stack | NxtGen AI · IOCL | TIET ’26

Role
Ai Engineer at Nxtgen Cloud Technologies Pvt Ltd
Location
Patiala, PB, IN
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Harsh Kumar

I am an AI/ML Engineer and Full-Stack Developer focused on building intelligent systems that go beyond prototypes into real production environments.My core work spans Agentic AI, Large Language Models, and Retrieval-Augmented Generation — designing systems capable of reasoning, multi-step planning, and autonomous task execution. I build production-grade RAG pipelines with hybrid retrieval, reranking, and query decomposition using LangChain and LlamaIndex, and architect stateful multi-agent workflows with LangGraph for enterprise automation.At NxtGen AI, I fine-tune open-source LLMs including LLaMA 3, Mistral, and NVIDIA Nemotron using LoRA/QLoRA for domain-specific tasks, and build real-time voice AI agents using Whisper ASR and neural TTS via LiveKit. I ship AI microservices with FastAPI and Docker in live cloud environments.Previously at Indian Oil Corporation Limited (IOCL), I built XGBoost-based predictive models for refinery operations achieving R² = 0.89, reducing manual lab dependency by 60% — results presented to and recognized by executive leadership.Beyond AI, I bring strong full-stack development experience with React, Next.js, Node.js, Flask, and Django, allowing me to build and deploy end-to-end AI-powered products — not just models.My final year project, Project ASTRA, is a multimodal smart assistant running on Raspberry Pi that combines RAG, gesture recognition, speech I/O, and IoT integration into a unified edge AI system.Core Skills- Agentic AI: LangChain, LangGraph, LlamaIndex, multi-agent systems, tool-calling, memory- LLMs: Fine-tuning (LoRA/QLoRA), RAG, prompt engineering, vector databases, embeddings- Voice AI: Whisper ASR, TTS, LiveKit- Models: LLaMA 3, Mistral, NVIDIA Nemotron, Gemini, Gptoss- Full-Stack: React, Next.js, Node.js, Flask, Django, FastAPI, Streamlit- Tools: Docker, Git, Firebase, SQL, Raspberry PiI am driven by a simple belief — AI should ship, not just impress. Always open to meaningful collaborations, research discussions, and opportunities at the intersection of AI and real-world engineering.

Experience

  1. Ai Engineer

    Nxtgen Cloud Technologies Pvt Ltd

    Jan 2026 — Present · Bengaluru, IN

    Building production-grade AI systems in the cloud infrastructure space — here\'s what I\'ve been working on: Architected end-to-end RAG pipelines using LangChain & LlamaIndex — hybrid retrieval, reranking, and query decomposition — cutting hallucination rates and boosting answer relevance by 40% Built stateful multi-agent systems with LangGraph featuring tool-calling, persistent memory, and self-correction loops for real enterprise automation workflows Fine-tuned LLaMA 3, Mistral & NVIDIA Nemotron using LoRA/QLoRA on domain-specific datasets; deployed on-premise with GGUF/AWQ quantization for optimized inference Engineered real-time voice AI agents — Whisper ASR + neural TTS via LiveKit — achieving sub-500ms end-to-end response latency for live conversational interfaces Shipped AI microservices with Docker + FastAPI, streaming LLM outputs to frontend clients via SSE in a production cloud environment

Education

  • Thapar Institute of Engineering & Technology

    Bachelor of Engineering, Computer Engineering

    2022 — 2026

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Harsh Kumar — Ai Engineer at Nxtgen Cloud Technologies Pvt Ltd in Patiala, PB, IN | Unifers