Rishabh Manoj
TPU Expert
- Role
- Senior Machine Learning Engineer at Google
- Location
- Hyderabad, TG, IN
- LinkedIn followers
- 500 followers
About Rishabh Manoj
As a Senior ML Engineer, I specialize in designing and developing customized machine learning and deep learning solutions that drive significant business impact by enhancing efficiency, optimizing costs, and generating revenue growth through AI.Currently, I am working at Google on cutting-edge SuperComputers, managing and optimizing clusters of thousands of GPUs and TPUs for large-scale LLM training. My focus is on capturing and analyzing telemetry data across these vast GPU and TPU clusters to provide insights for clients, ensuring peak performance and reliability in high-stakes environments.Key Achievements: Cost Savings: Saved ~$3M annually for a market research client with an NLP deep learning solution. Efficiency Improvement: Boosted user efficiency by over 70% with an NLP-based document management suite. Project Delivery: Delivered 30+ PoCs and 10 MVPs across Generative AI, Computer Vision, NLP, and Predictive Analytics. Team Leadership: Led a team of 10 ML engineers, providing expert guidance and mentorship. Technical Innovation: Created an LLM-based code generation platform, winning the Gold Award for \"Technical Innovation of the Year.\"My work also spans diverse LLM applications, including LLM-powered search engines, Text-to-SQL, and various multimodal solutions. I hold an Integrated M.Tech in Information Technology from where I built a solid foundation in machine learning, deep learning, computer vision, NLP, and data science.My goal is to leverage my skills and expertise to help clients transform their businesses through AI.
Experience
Senior Machine Learning Engineer
Aug 2024 — Present · Hyderabad, IN
Oversee and optimize thousands of GPU and TPU clusters across Google SuperComputers to support large-scale LLM training for internal and external clients, ensuring robust performance in complex, high-demand environments. Develop and maintain telemetry systems to capture real-time health, performance, and utilization metrics of GPU and TPU clusters, enabling proactive monitoring and diagnostics for enhanced system reliability. Collaborate closely with large clients to ensure uninterrupted workloads, implementing failover mechanisms and rapid issue resolution processes to mitigate the impact of hardware failures. Utilize telemetry insights to guide decisions on resource allocation, predictive maintenance, and performance tuning, resulting in optimal hardware utilization and significant efficiency gains across clusters. Drive cost optimization by analyzing and reducing system inefficiencies, balancing high-performance needs with sustainable resource management, ultimately lowering infrastructure expenses for clients. Leverage cutting-edge ML and observability tools to address complex infrastructure challenges, developing custom solutions that enhance system resilience and adaptability under heavy LLM training loads.
Education
International Institute of Information Technology Bangalore
Integrated Mtech, Information Technology
2013 — 2018
Skills
- Mysql
- Data Structures
- Algorithms
- Html5
- Python
- Javascript
- Php
- Html
- Android Development
- C++
- Java
- Sql
- Linux
- C
- Programming
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.