Arnav Kumar Aditya

AI-ML Engineer | Agentic AI | Generative AI | LLM | Machine Learning | Data Science | Amazon ML Summer School 2021

Role
Ai-ml Engineer at Cyfuture
Location
New Delhi, DL, IN
LinkedIn followers
500 followers

About Arnav Kumar Aditya

AI/ML Engineer with 1.5+ years of experience delivering production-ready solutions in Machine Learning, Generative AI, and Agentic AI systems. Skilled in designing scalable pipelines, fine-tuning LLMs, building RAG architectures, and deploying high-performance AI agents for real-world enterprise use cases.Hands-on expertise with LLaMA, Mistral, Whisper, GPT-OSS, and multimodal models, along with advanced deployment stacks including vLLM, llama.cpp, FastAPI, Docker, and Kubernetes. Experienced in GPU optimization, quantization, batching, and accelerating inference across cloud and on-prem GPU clusters.Adept at building autonomous AI agents using LangGraph, FastMCP, PGVector, and A2A protocols—enabling tool-calling, workflow automation, API integrations, and intelligent knowledge retrieval. Developed BYOM platforms, vector-search pipelines, and fine-tuned domain models for domain-specific tasks.Strong grounding in ML algorithms, deep learning, NLP, analytics, and data engineering with Python, SQL, Pandas, and Power BI. Focused on building reliable, impactful, and scalable AI systems that improve efficiency, automate workflows, and deliver measurable business value.What I Can Do & Collaborate On : • Build scalable AI systems, including LLM agents, RAG workflows, multimodal models, and voice-based applications.• Develop intelligent AI agents capable of tool use, automation, reasoning, and real-time decision-making.• Optimize and deploy machine learning models for low latency, high throughput, and reliable production performance.• Create end-to-end conversational and voice AI systems with speech recognition, synthesis, telephony integration, and real-time interaction.• Build knowledge retrieval pipelines, vector search systems, and document intelligence solutions.• Design production-ready ML infrastructure with containerization, orchestration, monitoring, and automated deployment.• Work with autonomous, agentic AI frameworks and multi-agent systems for complex workflows.• Implement safety-aligned, auditable, and controllable AI systems, and explore fine-tuning, model compression, and efficient deployment strategies.• Engage in AI research, experimentation, responsible AI practices, and emerging trends like LLM orchestration, personalized AI, and synthetic data generation.

Experience

  1. Ai-ml Engineer

    Cyfuture

    May 2025 — Present · Noida, IN

    Developed a production-grade GPT-OSS (20B) Voice Call Agent integrating Whisper ASR, Asterisk PBX, EdgeTTS, and RAG-based context retrieval, achieving <3s latency and 95% response accuracy in live customer conversations.• Built and deployed an Omni-Channel AI Chatbot across Web, WhatsApp, and Telegram, leveraging LangGraph, FastMCP, PGVector, and SearXNG for tool-calling, document retrieval, and real-time web search, serving monthly users.• Designed and engineered a BYOM (Bring Your Own Model) Platform enabling seamless deployment of LLMs, Embedding Models, ASR/TTS Models, and HuggingFace models, supporting 10+ models and 70+ daily inference requests with monitoring, scaling, and version control.• Optimized deployment pipelines for text, vision, and speech LLMs including GPT-OSS, LLaMA, Mistral, Whisper, achieving up to 3× throughput gains on AMD GPU clusters and NVIDIA H100 nodes through batching, quantization, and kernel-level tuning.• Built an LLM Finetuning Platform supporting Full Finetuning, LoRA, QLoRA, PEFT, dataset uploads, training configuration, and experiment tracking; fine-tuned LLaMA models (3B & 8B) achieving 20–30% higher task accuracy on customer datasets.• Architected scalable RAG workflows with enterprise-grade retrieval, vector search, and tool-calling using LangGraph + FastMCP, enabling agents to interact with APIs, knowledge bases, and search engines autonomously.• Deployed and maintained high-availability AI/ML services using FastAPI, Docker, KServe, Kubernetes, ensuring stable LLM inference, autoscaling, observability, and seamless integration with production workloads.• Enhanced GPU utilization and reduced inference cost through model quantization, caching, multi-model serving, and vLLM-based optimization.

Education

  • Kendriya vidyalaya

    10th, 90.20 %

    2018 — 2020

  • Maharaja Surajmal Institute Of Technology

    Bachelor of Technology - BTech, Computer Science, 9.396

    2020 — 2024

  • Kendriya Vidyalaya (KV)

    12th, Science, 90.80 %

    2019 — 2020

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Arnav Kumar Aditya — Ai-ml Engineer at Cyfuture in New Delhi, DL, IN | Unifers