Vishal Deep Verma
AI Voice & Video Agents for Real-World Use | Realtime Calls, Intelligent Routing, Zero-Downtime Systems
- Role
- Principal Ai Engineer Ai Voice, Video & Realtime Systems at Iwish
- Location
- Lucknow, UP, IN
- LinkedIn followers
- 500 followers
About Vishal Deep Verma
AI Voice & Video Agents only matter if they work in real-world conditions.I build production-grade systems that handle realtime calls, process speech (STT/TTS), and route conversations intelligently — with reliability as a core requirement, not an afterthought.My focus is on:Realtime voice & video infrastructureIntelligent call routing and automationLow-latency system designZero-downtime, fault-tolerant architecturesThis means designing for edge cases, handling failures gracefully, and ensuring systems perform consistently under load.Tech stack:Languages: Node.js, TypeScript, PythonAI/LLMs: OpenAI, LangChain, vector databasesRealtime & Voice: WebRTC, Twilio, streaming pipelinesBackend & Infra: AWS, Docker, Redis, message queuesInterested in building AI systems that operate beyond demos — where performance, scale, and reliability actually matter.
Experience
Principal Ai Engineer Ai Voice, Video & Realtime Systems
Dec 2024 — Present
Built and delivered production-grade AI voice systems with realistic concurrency, low latency, and scalable architecture.Product 1: AI Contact Center Platform (Multi-Industry, US Healthcare)Supports 200–800 concurrent calls per cluster (horizontal autoscaling)Latency:~500–1200 ms end-to-end (streaming STT → LLM → TTS)First response time:~300–600 ms (streaming partial responses)Multi-tenant omnichannel (voice/SMS/video/chat) with strict isolationReal-time voice pipeline: LiveKit + Deepgram + Twilio + Silero VADRAG pipelines (pgvector + embeddings) for tenant-scoped knowledgeYAML-driven workflow engine (routing, escalation, scheduling)CRM integrations via abstraction layer (Salesforce, Dynamics, Zoho, Vtiger)Backend: FastAPI + PostgreSQL (async, high I/O throughput)Agent orchestration: AutoGen / LLM tool-calling (GPT-4o class models)Infra: Dockerized microservices, queue-based load leveling, autoscalingTech: Python, FastAPI, PostgreSQL, pgvector, LLM tool-calling, LiveKit, Deepgram, Twilio, Docker, ReactProduct 2: Dylan AI (Automotive – Real-Time Voice AI)Handles 150–500 concurrent live voice sessions (scales via worker nodes)Latency:~400–900 ms streaming response timeCall setup:~1–2 sec (SIP/WebRTC)Built real-time streaming pipeline (LiveKit + Pipecat)Telephony: SIP, RTP, WebRTC with multi-provider abstractionMulti-agent orchestration, dynamic routing, mid-call handoffsRAG + prompt pipelines for grounded responsesReal-time transcription, summaries, sentiment, analyticsMicroservices: Python, Node.js, TypeScriptDeployment: AWS/Azure with autoscaling, health checks, failoverTech: Python, Node.js, TypeScript, LiveKit, Pipecat, Deepgram, WebRTC, LangChain, RAG, Docker, AWS, Azure
Education
Dr. A.P.J. Abdul Kalam Technical University
Bachelor's degree, Computer Programming
Skills
- Recruiting
- Business Strategy
- Strategy
- Resume Search
- IT Recruitment
- Team Management
- Microsoft Office
- Microsoft Word
- Powerpoint
- Microsoft Excel
- Sourcing
- Vendor Management
- Global Talent Acquisition
- Nxcam
- Microsoft Powerpoint
- Screening
- Autocad
- Catia
- Market Research
- Microsoft Outlook
- Leadership
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.