Abdul Rasheed
LLM Inference Performance Architect @ Qualcomm
- Role
- Staff Engineer at 高通
- Location
- San Diego, CA, US
- LinkedIn followers
- 500 followers
About Abdul Rasheed
Staff Engineer working on large‑scale LLM inference systems and performance optimization for AI accelerators.My work sits at the intersection of LLM architecture, compiler/runtime systems, and hardware acceleration, with a focus on latency‑critical and throughput‑efficient inference. I’ve led and contributed to end‑to‑end optimizations across model execution, pipeline parallelism, KV‑cache management, disaggregated serving, and HW‑aware kernel execution on production AI inference platforms.I enjoy turning complex models into measurable performance gains — reducing TTFT, improving users/min, and enabling scalable serving for large foundation models.Interests: LLM inference at scale, HW/SW co‑design, performance modeling, distributed & disaggregated serving, and next‑gen AI accelerators.
Experience
Staff Engineer
Nov 2025 — Present · San Diego, CA, US
Education
University of Connecticut
Masters, Electrical Computer Engineering
Forman Christian College (A Chartered University)
HSSC, Pre-Engineering
University of Engineering and Technology, Lahore
Bachelor's degree, Electrical and Electronics Engineering
Skills
- C
- Programming
- Java
- Matlab
- Embedded C
- Multithreading
- Verilog
- Labview
- Microsoft Word
- Linux
- Computer Vision
- Windows
- Assembly Language
- Real-Time Operating Systems (Rtos)
- Testing
- Embedded Systems
- Nucleus Rtos
- Microsoft Office
- C++
- Microsoft Excel
- Vhdl
- Simulations
- Microsoft Powerpoint
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.