Thristha Gurajala

AI/ML Engineer | LLM & RAG Systems | Multimodal AI | PyTorch | AWS & Kubernetes | Distributed Inference | Responsible AI

Role
Ai Ml Engineer at Perplexity
Location
San Francisco, CA, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Thristha Gurajala

I’m an AI/ML Engineer with 5+ years of experience designing, deploying, and optimizing…

Experience

  1. Ai Ml Engineer

    Perplexity

    May 2025 — Present · San Francisco, CA, US

    Reduced average inference latency by 42% through hybrid orchestration of on-device quantized models and AWSGPU inference clusters using speculative decoding, cutting end-user wait time from 2.6 s → 1.5 s.• Increased AI task completion accuracy by 31% via reinforcement-driven retrieval-augmented generation (RAG)pipeline and fine-tuned Sonar LLMs with context-aware embeddings.• Improved user engagement by 27% MoM by launching personalization and memory layers using encrypted localstorage + vector similarity ranking (FAISS on AWS EKS), driving measurable user retention lift.• Designed and deployed distributed inference architecture on AWS EKS leveraging PyTorch 2.x, Triton InferenceServer, and Ray Serve, with autoscaling and GPU node pool segregation for 10K+ concurrent sessions.• Engineered real-time media processing pipelines with Kafka + Flink, FFmpeg, and Whisper STT to enable multimodalsummarization (video/audio/text) inside the Comet Browser assistant.• Implemented secure browser-side AI agents using TypeScript + WebGPU + WASM for low-latency summarizationand form-action inference directly within Chromium render processes.• Built and optimized retrieval layer using FAISS + Elasticsearch for 20 B+ documents, enabling sub-200 ms vectorlookup times and citation-aware contextual grounding.• Automated ML ops pipelines via AWS SageMaker, Kubernetes, and GitHub Actions, integrating model registry, CI/CD,and A/B testing across production Sonar LLMs.• Integrated AdTech and payment microservices (Stripe API, AWS Lambda, DynamoDB) with revenue-share logic forComet Plus subscriptions and publisher monetization dashboards.• Deployed advanced observability stack (Prometheus, Grafana, OpenTelemetry, Sentry) with real-time drift detectionand prompt-injection monitoring, improving incident resolution speed by ~45%.

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Thristha Gurajala — Ai Ml Engineer at Perplexity in San Francisco, CA, US | Unifers