Rohit Kushwaha
AI Engineer | Building Low-Latency Enterprise GenAI Products | RAG, LLM Fine-Tuning, Multi-Agent Systems
- Role
- Artificial Intelligence Engineer at W3villa Technologies
- Location
- Kanpur Nagar, UP, IN
- LinkedIn followers
- 500 followers
About Rohit Kushwaha
I bridge the gap between research-grade AI and production-ready applications. As a AI Engineer, I don’t just train models; I build the infrastructure that makes them fast, reliable, and profitable.CORE IMPACT: • Scale: Engineered systems handling 150k+ high-volume transactions and 120k+ monthly AI interactions. • Performance: Cut query latency by 40% using TensorRT quantization and vector store optimization. • Innovation: Built privacy-first offline voice tutors and complex multi-agent orchestration (OpenAI Swarm).TECHNICAL TOOLKIT: • AI/ML: LangChain, RAG (FAISS/Pinecone), Fine-tuning (LoRA/QLoRA), PyTorch, Hugging Face. • Architecture: Microservices, Agentic Workflows, CI/CD, Docker. • Frontend/Backend: Python, ReactJS, FastAPI.THE \"LAST MILE\" PHILOSOPHY: Most AI fails in production because of latency and hallucination. I specialize in solving these through advanced RAG architectures and quantization for edge deployment.Let’s build the future of GenAI. Reach out for: AI Architecture Strategy, Agentic Workflows, or Production-grade RAG. Based in India | Open to global collaboration.
Experience
Artificial Intelligence Engineer
Feb 2023 — Present
Orchestrated multi-tenant AI agents for Slack and WhatsApp (HRMS, CRM, LMS) using Python, Google ADK, and MCP, managing 120k+ monthly interactions with 98% intent accuracy. Enhanced agent reliability by fine-tuning Gemma-7B and Function Gemma models using LoRA and PEFT, achieving higher determinism in complex tool-calling scenarios. Architected a high-performance RAG pipeline utilizing LlamaIndex, FAISS, and Neo4j Knowledge Graphs to enable context-aware search across 50k+ enterprise records, reducing query latency by 35%. Developed a vendor-agnostic LLM Proxy Gateway supporting OpenAI, Gemini, Mistral, and Groq with built-in RBAC and intelligent routing. Implemented KV-cache and semantic prompt caching mechanisms to optimize token usage, significantly reducing inference costs and improving latency under high concurrency. Boosted platform scalability by migrating from WebSockets to Server-Sent Events (SSE) and implementing table partitioning and sharding, leading to substantial gains in query throughput. Engineered a production-grade CDC pipeline from PostgreSQL to Elasticsearch using logical replication and WAL decoding for real-time, zero-drift data synchronization. Established enterprise-grade LLM safety by implementing dynamic guardrails and policy validation layers to prevent prompt injection and data leakage. Designed a distributed Redis Sliding Window Rate Limiter to protect critical AI endpoints from traffic spikes and abuse. Built a fault-tolerant, Celery-based distributed task queue for chat processing with automatic retries and centralized monitoring. Achieved 99.9% uptime for multi-tenant AI workloads through Docker-based containerization, resource profiling, and blue-green deployment strategies. Integrated end-to-end encrypted messaging with client-side cryptography to ensure a zero-knowledge architecture and maximum user privacy.
Education
Kanpur Institute of Technology
Master of Computer Applications - MCA, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.