Shreyashri Biswas
Deep Learning Engineer @AMD
Signup · Get unlimited contacts
WORK HISTORY
Deep Learning Engineer @AMD
Austin, TX, US
Accelerating AI at AMD through full stack optimization across ROCm- Engineered high-performance FBGEMM GPU kernels- Architected and deployed a distributed inference pipeline for DeepEP and rocSHEMM- Designed custom Triton StreamK GEMM kernels with major throughput gains- Performance tuning for LLMs like Llama, DeepSeek, and Mistral across PyTorch, vLLM, and RCCL- Optimized llama.cpp with advanced quantization and memory strategies- Supported CuPy enablement for ROCm by stabilizing GPU backend and resolving critical regressions- Refined compiler IR and assembly to unlock substantial speedups in AI workloads.
EDUCATION
Carnegie Mellon University
Master of Science - MS, Electrical and Computer Engineering
SRM IST Chennai
Bachelor of Technology - BTech, Electrical, Electronics and Communications Engineering
ABOUT SHREYASHRI BISWAS
HW SW co design for scalable inference pipelines Disaggregated prefill decode and end…
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.