Rong Ou

Software Engineer AI/ML/Data

Role
Principal Systems Software Engineer at NVIDIA
Location
Palo Alto, CA, US
LinkedIn followers
500 followers
Information TechnologyView LinkedIn profile

About Rong Ou

Technical leader and software engineer steeped in the confluence of large-scale machine learning, massively parallel and distributed systems and infrastructure, and good software engineering practices. Interested in solving challenging problems in AI/ML with high impact.

Experience

  1. Principal Systems Software Engineer

    NVIDIA

    Jun 2017 — Present · Santa Clara, CA, US

    Accelerated Machine Learning and Big Data | 2018/11 - present* Implemented the arena memory allocator in RAPIDS Memory Manager (RMM) that improved GPU memory allocation speed, and reduced fragmentation in a highly concurrent environment typical of Apache Spark. It became the default allocator for Spark RAPIDS and unblocked per-thread default stream for fully asynchronous CUDA operations.* Implemented federated learning from scratch in XGBoost, for both horizontal/vertical settings and CPU/GPU based tree methods. Designed a new algorithm to use bit masks to speed up training and reduce communication overhead. Implemented MPI primitives (allreduce, allgather, broadcast) using gRPC. Federated XGBoost is being adopted by large financial institutions.* Implemented GPU out-of-core training in XGBoost. This involved a lot of cleanup and refactoring of the code base, and implementing gradient-based sampling. Enabled training much larger datasets on a given GPU, without degrading model accuracy or training time.* Added GPUDirect Storage support to Spark RAPIDS, improved the spilling algorithm to better support processing large datasets.* Helped ramp up the Spark RAPIDS team in Shanghai. Wrote a design doc for encoding categorical features which was successfully implemented by the team.* Wrote the design proposal for federated learning for medical imaging, which led to the founding of the NVFlare project and several patents.AI Infrastructure for Autonomous Vehicles | 2017/06 - 2018/10* Lead architect for a machine learning platform built on top of Kubernetes, with the goal to make it easy to do large-scale deep learning in a production environment.* Implemented the Kubernetes MPI Operator to manage multi-node distributed training jobs optimized for GPUs. The project is open sourced as part of Kubeflow.* Researched into self-supervised learning and reinforcement learning for self-driving cars, leveraging the large amount of unlabeled videos and sensor data.

Education

  • Peking University

    Bachelor of Arts (B.A.), Physics

    1988 — 1992

  • The University of Texas at Austin

    Master of Science (M.S.), Computer Science

    1994 — 1996

Skills

  • Java Enterprise Edition
  • Oop
  • Tdd
  • Deep Learning
  • Software Development
  • Nosql
  • Python
  • Big Data
  • Object Oriented Design
  • Machine Learning
  • Android
  • Technical Leadership
  • Software Design
  • Agile Methodologies
  • Agile
  • Tensorflow
  • Java
  • Artificial Intelligence
  • Scala
  • Go
  • Representational State Transfer (Rest)
  • Extreme Programming
  • C++
  • Rest
  • Distributed Systems
  • Test Driven Development
  • Scalability
  • Software Engineering
  • Object-Oriented Programming (Oop)
  • Cloud Computing
  • Refactoring
  • Junit
  • Sql

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Rong Ou — Principal Systems Software Engineer at NVIDIA in Palo Alto, CA, US | Unifers