Vamsi Sripathi
Member of Technical Staff, Microsoft AI
- Role
- Member of Technical Staff at Microsoft
- Location
- Mountain View, CA, US
- LinkedIn followers
- 500 followers
About Vamsi Sripathi
15 years of experience in applying x86 code optimizations and thread parallelism to mathematical libraries, deep learning frameworks and scientific applications on multiple generations of Intel processors- 8 years of leadership experience in driving HPC/AI technical engagements with cross-org team ofjunior/mid-level staff engineers- 6 years of software product development experience in Intel Math Kernel Library- Demonstrated track-record of identifying performance bottlenecks and optimizations for compute and memory bound workloads through robust understanding of CPU and GPU micro-architecture- Contributions to Intel hardware/software co-design efforts through collaborations with lead hardwareand Compiler architects in the evaluation of Intel hardware features (Data Streaming Accelerator, Spec-ulative i2m) and ISA extensions (avx512, Vector Neural Network Instructions (VNNI)).• Key accomplishments- Directly contributed to several Intel Silicon design wins valued at $M’s by delivering targeted code optimizations- Upstreamed code optimizations (avx512 vectorization, prefetching) to deep learning frameworks (TensorFlow, Caffe, Eigen) and HPC domains (Climate/Weather, QCD, Geospatial). Optimized avx512 instruction sequence for prefix-sum, argmax that beats Intel, GCC, Clang Compiler performance- Enabled key external customers and collaborated with cross-organizational teams to enhance the positioning of Intel platforms in the HPC/AI markets. Worked with a wide spectrum of customers - US/Europe National Labs (ORNL, ANL, NCAR, ECMWF, TUDA), CSPs (Meta, Amazon), Taboola, Lenovo, HPE, GE, Siemens, Ford, General Motors- Synthesized complex HPC/machine learning workloads to representative compute/memory bound kernels used in performance debug (emulation, post-Silicon) of Intel CPUs and GPUs- Implemented avx, avx2, avx512 optimizations to Basic Linear Algebra Subroutines (BLAS)/matrix operations in Intel MKL. Designed and developed compiler SIMD vector intrinsics framework for MKL BLAS optimizations. Robust product development experience spanning 5 major releases of MKL- Performance optimization of MPI workloads on large supercomputers, improved MPI I/O performance on Lustre file-system at 100k processes of Cray XT supercomputer - Senior Member of ACM and Contributions to Intel platform tuning guides, publications in ACM conferences
Experience
Member of Technical Staff
Aug 2025 — Present · Mountain View, CA, US
Education
North Carolina State University
M.S., Computer Science
2007 — 2010
Skills
- Performance Analysis
- Linux
- Intel Architecture
- Algorithms
- Latex
- Multi-Core
- Parallel Computing
- Scalability
- Parallel Programming
- Simd
- Bash
- Hdf5
- Optimizing Performance
- Mpi
- Computer Architecture
- Distributed Systems
- X86 Assembly
- C
- Optimization
- Openmp
- Lustre
- High Performance Computing
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.