Robert McQueen

Senior Ai Ml Hpc Cluster Engineer @NVIDIA

Peoria, AZ, US
MOBILE NUMBERS
+16•••••••22

Signup · Get unlimited contacts

WORK HISTORY

May 2024 — Present

Senior Ai Ml Hpc Cluster Engineer @NVIDIA

View department →

Peoria, AZ, US

Lead operations of large-scale AI/HPC clusters across on-premises and cloud environments, serving as primary “pilot in command” for upgrades, reliability improvements, and incident response. • Drive end-to-end cluster lifecycle management, including deployment of heterogeneous compute, networking (InfiniBand/RDMA), and high-performance storage (Lustre, GPFS). • Designed and implemented scalable automation frameworks leveraging Ansible, Kubernetes, Docker/Singularity, and Python, improving system reliability and reducing manual intervention. • Optimized GPU-accelerated workloads through performance tuning, fragmentation reduction, and resource utilization analysis, cutting GPU waste and increasing throughput against SLA targets. • Partner with researchers and ML engineers to analyze and optimize deep learning workflows (MPI, CUDA, PyTorch, TensorFlow), enabling faster experimentation and reduced model training time.

SKILLS

HardwareVeritas Cluster ServerUnix SoftwareHp OpenviewVeritas Volume ManagerStorageComputer HardwareDisaster RecoveryData CenterUnixVendor ManagementVirtualizationSolarisTroubleshootingWindowsSecurityIntegrationLinuxDocumentationArchitectureXpServer ArchitectureAutomationEngineeringHpTestingAdministrationServersProblem SolvingStorage Area NetworksTcp/IpOperations ManagementHp-UXCustomer RelationsSoftware InstallationVendor RelationsRequirements AnalysisSoftware DevelopmentAnalysisVmware

ABOUT ROBERT MCQUEEN

Highly skilled, technically oriented professional with more than nineteen years…

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.