Steffen Hirschmann
Senior Software Performance Engineer, Co-founder @Efficientware
Signup · Get unlimited contacts
WORK HISTORY
Senior Software Performance Engineer, Co-founder @Efficientware
Stuttgart, DE
Boosted vLLM inference throughput by 15% on an 8× NVIDIA H200 cluster with a custom MoE monokernel implementation in CUDA that doubled MoE throughput in vLLM.→ Accelerated person detection CNN inference based on TensorFlow-lite micro on Cortex-M4 from 190 ms to 95 ms (2× latency reduction/throughput increase) by optimizing ARM CMSIS-NN kernels; presented 2× speed-up at Embedded World 2025.→ Drove the vision to position Efficientware as edge-AI performance optimization company; we successfully scaled these services to data-center GPU clusters.
EDUCATION
University of Stuttgart
Doctor of Science (PhD)
ABOUT STEFFEN HIRSCHMANN
Performance profiling & diagnostics: Linux Perf, Intel VTune, NVIDIA Nsight, custom instrumentation→ Runtime optimization: Latency and throughput tuning on Intel, AMD, ARM CPUs, NVIDIA GPUs,Texas Instruments DSPs, Qualcomm Hexagon NPUs→ Cross-platform porting: Migrating kernels & applications from CPU to GPU, DSP, and NPU→ Technical mentoring: Trained 25+ peers in 3+ workshops→ Project experiences:>1M LoC, people (advanced driver assistance systems)→ Languages: C, modern C++, CUDA, Python, Assembly→ Technologies: MPI, OpenMP, AVX, NEON; PyTorch, vLLM; ARM CMSIS-NN, TF lite-micro
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.