Zhang Jiejing
Principal Engineer @Together AI
Signup · Get unlimited contacts
WORK HISTORY
Principal Engineer @Together AI
San Francisco, CA, US
Led end-to-end LLM inference acceleration for production models and key customers, covering profiling → bottleneck isolation → optimization design → implementation → validation.Drove performance improvements across core kernels and serving stack: GEMM, Attention/FlashAttention-style optimizations, MoE operators, CUDA graph/launch overhead, and memory-bandwidth bottlenecks.Designed and implemented cache-aware disaggregated inference (Prefill/Decode separation), enabling scalable throughput under high concurrency and long-context workloads.Built/optimized distributed KV-Cache architecture and cache management strategies (cache locality, paging/eviction, network + GPU memory tradeoffs) to reduce tail latency and improve cluster efficiency.Established performance methodology and tooling (GPU profiling, roofline-style analysis, kernel-level metrics, end-to-end QPS/TTFT/TPOT) and delivered actionable optimization plans for multiple model families.
EDUCATION
University of Electronic Science and Technology of China
Bachelor, Software, Computer Science and technology
SKILLS
ABOUT ZHANG JIEJING
A wild areas programmer, includes Linux Kernel, Linux Kernel Driver,Linux file system, Linux memory management, GNU Tool-Chain, GUI programming, Linux Application Programmer, Network and script programmer.Specialties: Write and read source code.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.