Robert McQueen
Senior Ai Ml Hpc Cluster Engineer @NVIDIA
Signup · Get unlimited contacts
WORK HISTORY
Senior Ai Ml Hpc Cluster Engineer @NVIDIA
Peoria, AZ, US
Lead operations of large-scale AI/HPC clusters across on-premises and cloud environments, serving as primary “pilot in command” for upgrades, reliability improvements, and incident response. • Drive end-to-end cluster lifecycle management, including deployment of heterogeneous compute, networking (InfiniBand/RDMA), and high-performance storage (Lustre, GPFS). • Designed and implemented scalable automation frameworks leveraging Ansible, Kubernetes, Docker/Singularity, and Python, improving system reliability and reducing manual intervention. • Optimized GPU-accelerated workloads through performance tuning, fragmentation reduction, and resource utilization analysis, cutting GPU waste and increasing throughput against SLA targets. • Partner with researchers and ML engineers to analyze and optimize deep learning workflows (MPI, CUDA, PyTorch, TensorFlow), enabling faster experimentation and reduced model training time.
SKILLS
ABOUT ROBERT MCQUEEN
Highly skilled, technically oriented professional with more than nineteen years…
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.