Rong Ou
Software Engineer AI/ML/Data
- Role
- Principal Systems Software Engineer at NVIDIA
- Location
- Palo Alto, CA, US
- LinkedIn followers
- 500 followers
About Rong Ou
Technical leader and software engineer steeped in the confluence of large-scale machine learning, massively parallel and distributed systems and infrastructure, and good software engineering practices. Interested in solving challenging problems in AI/ML with high impact.
Experience
Principal Systems Software Engineer
Jun 2017 — Present · Santa Clara, CA, US
Accelerated Machine Learning and Big Data | 2018/11 - present* Implemented the arena memory allocator in RAPIDS Memory Manager (RMM) that improved GPU memory allocation speed, and reduced fragmentation in a highly concurrent environment typical of Apache Spark. It became the default allocator for Spark RAPIDS and unblocked per-thread default stream for fully asynchronous CUDA operations.* Implemented federated learning from scratch in XGBoost, for both horizontal/vertical settings and CPU/GPU based tree methods. Designed a new algorithm to use bit masks to speed up training and reduce communication overhead. Implemented MPI primitives (allreduce, allgather, broadcast) using gRPC. Federated XGBoost is being adopted by large financial institutions.* Implemented GPU out-of-core training in XGBoost. This involved a lot of cleanup and refactoring of the code base, and implementing gradient-based sampling. Enabled training much larger datasets on a given GPU, without degrading model accuracy or training time.* Added GPUDirect Storage support to Spark RAPIDS, improved the spilling algorithm to better support processing large datasets.* Helped ramp up the Spark RAPIDS team in Shanghai. Wrote a design doc for encoding categorical features which was successfully implemented by the team.* Wrote the design proposal for federated learning for medical imaging, which led to the founding of the NVFlare project and several patents.AI Infrastructure for Autonomous Vehicles | 2017/06 - 2018/10* Lead architect for a machine learning platform built on top of Kubernetes, with the goal to make it easy to do large-scale deep learning in a production environment.* Implemented the Kubernetes MPI Operator to manage multi-node distributed training jobs optimized for GPUs. The project is open sourced as part of Kubeflow.* Researched into self-supervised learning and reinforcement learning for self-driving cars, leveraging the large amount of unlabeled videos and sensor data.
Education
Peking University
Bachelor of Arts (B.A.), Physics
1988 — 1992
The University of Texas at Austin
Master of Science (M.S.), Computer Science
1994 — 1996
Skills
- Java Enterprise Edition
- Oop
- Tdd
- Deep Learning
- Software Development
- Nosql
- Python
- Big Data
- Object Oriented Design
- Machine Learning
- Android
- Technical Leadership
- Software Design
- Agile Methodologies
- Agile
- Tensorflow
- Java
- Artificial Intelligence
- Scala
- Go
- Representational State Transfer (Rest)
- Extreme Programming
- C++
- Rest
- Distributed Systems
- Test Driven Development
- Scalability
- Software Engineering
- Object-Oriented Programming (Oop)
- Cloud Computing
- Refactoring
- Junit
- Sql
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.