Surendhar Chowdary
Senior AI/ML Engineer | Generative AI & LLM Systems | RAG, NLP, Transformers | Cloud-Native (GCP) | Building Scalable, Production-Grade ML Platforms
- Role
- Ai Ml Engineer at Meta
- Location
- San Francisco, CA, US
- LinkedIn followers
- 500 followers
About Surendhar Chowdary
AI/ML Engineer with 5+ years of experience delivering high-impact machine learning and Generative AI solutions across enterprise environments. Strong foundation in core ML, NLP, and predictive modeling, with advanced expertise in building production-grade LLM systems using RAG pipelines, embeddings, and transformer architectures. Proven track record of designing end-to-end scalable systems from data processing to deployment leveraging LangChain, Hugging Face, and OpenAI APIs on GCP with GPU-enabled infrastructure. Skilled in FastAPI, Docker, Kubernetes, and MLOps (MLflow, Airflow), with a focus on optimizing performance, reducing hallucinations, and driving measurable business outcomes.
Experience
Ai Ml Engineer
Jun 2024 — Present · San Francisco, CA, US
Designed and deployed end-to-end Generative AI systems using LLMs and RAG pipelines, improving response accuracy by 30%+ across enterprise applications. Built scalable RAG pipelines using LangChain, Hugging Face, and FAISS/Pinecone for context-aware Q&A with reduced hallucinations. Developed and optimized transformer models (BERT, GPT variants) using PyTorch on GPU, reducing inference latency by 25%. Implemented embedding pipelines using OpenAI and Hugging Face for semantic search and recommendation systems on large datasets. Designed data ingestion and preprocessing pipelines for structured/unstructured data using Python, Pandas, and BigQuery (GCP). Built and deployed REST APIs with FastAPI to serve ML/LLM models for real-time inference and scalable integration. Leveraged GCP (Vertex AI, BigQuery, Cloud Storage, Compute Engine) for training, deployment, and large-scale processing. Containerized applications with Docker and deployed on Kubernetes (GKE) for high availability and fault tolerance. Implemented MLOps pipelines using MLflow and Airflow for experiment tracking, versioning, and workflow automation. Designed real-time pipelines using Kafka for streaming data and near real-time predictions. Performed LLM evaluation using TruLens and OpenAI Evals, improving reliability and reducing hallucinations. Applied prompt engineering and fine-tuning (LoRA/PEFT) for domain-specific LLM customization. Monitored models using logging, alerting, and drift detection to ensure consistent production performance. Collaborated with cross-functional teams to design AI architectures and deliver scalable ML solutions. Built NLP pipelines for semantic search, classification, and entity extraction, improving RAG performance.
Education
Auburn University at Montgomery
Master's Degree, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.