Nikolay Krapivnyy
ML infrastructure at Google
- Role
- Site Reliability Engineering Manager, Ml Compute Sre at Google
- Location
- London, GB
- LinkedIn followers
- 500 followers
About Nikolay Krapivnyy
Engineering leader with deep technical skills. Result-oriented, have great troubleshooting and analytical skills. Over 20 years in IT: from developer to SRE / Engineering Director. Experienced in:– building and running SRE teams with focus on defining and effectively defending SLO numbers for large scale Cloud and ML infrastructure– managing product delivery teams to achieve business goals via reducing time to market, faster iterations, agile development process– building and running multi-brand trust and safety function – introducing data-driven approach for product delivery pipeline
Experience
Site Reliability Engineering Manager, Ml Compute Sre
Jun 2024 — Present
Leading SRE team responsible for- Reliability of Google\'s ML fleet (both TPU and GPU accelerators)- Large scale training and inference infrastructure (LLMs)- Defining and defending SLO targets for critical large scale ML workloads- ML workloads scheduling logic and optimisation for TCO and utilization numbers improvements
Education
Moscow Power Engineering Institute (Technical University)
Master's degree, Computer Science
2005 — 2011
Skills
- Php
- Git
- Mysql
- Web Development
- Ajax
- Nginx
- Bash
- Oop
- Mvc
- Linux
- Apache
- Mobile Applications
- Jquery
- Html
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.