Nikolay Krapivnyy

Nikolay Krapivnyy

ML infrastructure at Google

Role
Site Reliability Engineering Manager, Ml Compute Sre at Google
Location
London, GB
LinkedIn followers
500 followers

About Nikolay Krapivnyy

Engineering leader with deep technical skills. Result-oriented, have great troubleshooting and analytical skills. Over 20 years in IT: from developer to SRE / Engineering Director. Experienced in:– building and running SRE teams with focus on defining and effectively defending SLO numbers for large scale Cloud and ML infrastructure– managing product delivery teams to achieve business goals via reducing time to market, faster iterations, agile development process– building and running multi-brand trust and safety function – introducing data-driven approach for product delivery pipeline

Experience

  1. Site Reliability Engineering Manager, Ml Compute Sre

    Google

    Jun 2024 — Present

    Leading SRE team responsible for- Reliability of Google\'s ML fleet (both TPU and GPU accelerators)- Large scale training and inference infrastructure (LLMs)- Defining and defending SLO targets for critical large scale ML workloads- ML workloads scheduling logic and optimisation for TCO and utilization numbers improvements

Education

  • Moscow Power Engineering Institute (Technical University)

    Master's degree, Computer Science

    2005 — 2011

Skills

  • Php
  • Git
  • Mysql
  • Web Development
  • Ajax
  • Nginx
  • Bash
  • Oop
  • Mvc
  • Linux
  • Apache
  • Mobile Applications
  • Jquery
  • Html

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Nikolay Krapivnyy — Site Reliability Engineering Manager, Ml Compute Sre at Google in London, GB | Unifers