Matthew Nikahd
Principal AI Infrastructure Engineer at Microsoft | Large-Scale GPU & Networking Systems | Researching High-Performance GPU Fabrics for Next-Gen LLMs
- Role
- Principal Ai Infrastructure Engineer at Microsoft
- Location
- San Jose, CA, US
- LinkedIn followers
- 500 followers
About Matthew Nikahd
I have over 15 years of experience in Systems and Networking, with a focus on both classic infrastructure and large-scale AI networking systems. My main interests lie in optimizing AI workflows by improving infrastructure-level components such as topology design, routing strategies, and congestion control mechanisms across both InfiniBand and RoCEv2 fabrics. I worked on RoCEv2 optimization (especially around low-entropy issues) and contributed to large-scale topology design, including the Rail-only topology. I also have a strong software engineering foundation and have built scalable automation and telemetry frameworks to improve operational efficiency and performance.In parallel, I’ve developed a strong understanding of deep learning and LLMs as problem domains, which helps me better align infrastructure decisions with ML workloads. I’m also engaged in related academic activities, including research and collaboration on applying machine learning to solve core networking challenges—particularly through network foundation models fine-tuned for tasks like traffic classification, anomaly detection, and burst prediction.#TCP/IP - Routing & Switching (BGP OSPF, Multicast Routing, switching technologies/ protocols)#Datacenter (Overlay BGP in Web-Scale Datacenter, BGP EVPN, VXLAN, SDN, L3 Fabric/Nexus)#InfiniBand, RoCE networks, RDMA, Congestion Control#Data structure, Advanced Algorithm Design#Software analysis and design#Python, C++#CI/CD, Pipeline#Predictive Modeling#Deep Learning#Diffusion Models#Generative & hl=en
Experience
Principal Ai Infrastructure Engineer
Sep 2023 — Present · San Francisco, CA, US
Leading the design and optimization of large-scale AI infrastructure networks, and working closely with providers such as CoreWeave, Nebius, Nscale and Lambda—covering over 300K GPUs powering Microsoft AI and OpenAI workloads.Additionally, conducting research focused on high-performance networking, scalability, and efficiency across multi-region GPU datacenters supporting next-generation LLM training and inference.
Education
Arizona State University
Master's degree, Computer Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.