Gayathri T
AI/ML Engineer @ Meta | Building Production-Scale LLM & RAG Systems | GenAI, Evaluation, MLOps | PyTorch • FastAPI • Node.js 🚀
- Role
- Ai Ml Engineer at Meta
- Location
- San Jose, CA, US
- LinkedIn followers
- 500 followers
About Gayathri T
AI/ML Engineer specializing in Generative AI, LLMs, and scalable ML systems, with 5+ years of experience delivering production-grade solutions in high-impact environments.I focus on building end-to-end AI systems that actually work in the real world — from fine-tuning and evaluating large language models to designing robust Retrieval-Augmented Generation (RAG) pipelines and deploying low-latency inference systems at scale. My experience spans optimizing model performance, reducing hallucinations, and ensuring reliability in enterprise-grade AI applications.With a strong foundation in MLOps and distributed systems, I’ve worked on high-throughput, real-time AI platforms, combining deep learning with efficient infrastructure to deliver measurable business impact. I’m particularly interested in solving challenges around LLM evaluation, system scalability, and production readiness.Currently exploring opportunities and collaborations in:• Generative AI / LLM Engineering• Applied AI & ML Systems• Scalable AI Infrastructure & MLOpsIf you’re working on cutting-edge AI problems or building impactful GenAI products, let’s connect.
Experience
Ai Ml Engineer
Mar 2024 — Present · US
Designed and deployed production-grade AI systems using Python and LLMs at Meta, enabling scalable, low-latency, multi-turn conversational experiences across internal AI platforms. • Fine-tuned LLaMA 3 based foundation models using supervised fine-tuning (SFT) and parameter-efficient techniques (LoRA/QLoRA), improving response grounding, instruction-following, and contextual reasoning. • Evaluated and benchmarked multiple foundation models including Mistral 7B and Mixtral to optimize trade-offs across latency, cost, and generation quality. • Built advanced Retrieval-Augmented Generation (RAG) pipelines using hybrid retrieval (dense and keyword), optimizing chunking strategies and embeddings to reduce hallucinations and improve factual grounding. • Developed semantic retrieval and vector search systems using Python and FAISS, incorporating scalable vector database design patterns inspired by Pinecone. • Leveraged Hugging Face Transformers for model experimentation, fine-tuning workflows, and tokenizer optimization; built modular pipelines using LangChain and LlamaIndex for orchestration and data integration. • Engineered high-performance inference systems in Python using PyTorch, optimizing latency and throughput via dynamic batching, KV caching, token streaming, and GPU utilization strategies across distributed infrastructure. • Deployed and scaled workloads on Meta’s internal cloud infrastructure, leveraging distributed GPU clusters and Tupperware-based container orchestration to support high-throughput, low-latency inference. • Developed scalable service APIs using Python (FastAPI), REST, and gRPC, enabling asynchronous processing and high-concurrency real-time generation workloads. • Contributed to distributed training pipelines for large models using PyTorch FSDP and DeepSpeed, applying mixed precision and memory optimization techniques to improve training efficiency and reduce compute cost.
Education
University of Central Missouri
Master's degree, Big Data Analytics & Information Technology
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.