Poorna Ravuri
Graduate Research Assistant @San José State University Research Foundation
Signup · Get unlimited contacts
WORK HISTORY
Graduate Research Assistant @San José State University Research Foundation
San Jose, CA, US
Optimizing LLM Inference with Quantization on Large-Scale AI Systems• Implemented a 5/6-bit quantization method using Bit-Sliced Indexing (BSI) and hybrid compression, reducing LLM memory footprint by 50% against an fp16 baseline with less than 6% accuracy loss.• Wrote custom multi-threaded C++/CUDA kernels for core tensor operations (matmul, dot product), exposing them to PyTorch via PyBind11 to achieve a 1.2x performance uplift on NVIDIA A100 GPUs.• Optimized kernel execution by using persistent-thread kernels and multi-stream concurrency to overlap data transfers and compute; accelerated CPU paths with AVX2/AVX-512 SIMD intrinsics.• Built a reusable, end-to-end LLM training pipeline on an multi-node GPU cluster using SLURM, automating CUDA compilation with CMake and CI scripts to reduce per-epoch training time.• Ensured code robustness and memory safety across the entire C++/CUDA codebase by passing rigorous Valgrind and ASAN checks with zero memory leaks.
EDUCATION
San José State University
Master's degree, Artificial Intelligence
GITAM Deemed University
Bachelor of Technology - BTech, Ece
ABOUT POORNA RAVURI
Intern at DigitalOcean | MS Artificial intelligence @ SJSU
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.