vLLM 0.28.0 Optimizes Open-Weight Inference
vLLM 0.28.0, released August 26, delivers a major Kimi K3 performance push alongside deeper DeepSeek V4 and Qwen3.8 support.
This release makes model architecture inseparable from serving economics: open-weight winners increasingly depend on the stack that turns exotic designs into affordable tokens. The gains remain hardware- and workload-dependent, so teams should benchmark their exact concurrency, context, quantization, and accelerator mix.
- –Kimi K3 gains fused FlashKDA kernels, decode-context parallelism, adaptive speculative decoding, and expert sharding that can save about 17 GiB per GPU.
- –vLLM reports up to 3.14× higher Kimi K3 throughput with DSpark on 16 NVIDIA GB300 GPUs. [Kimi K3 benchmarks](https://vllm-project.github.io/2026/07/27/k3.html)
- –DeepSeek V4 receives end-to-end sparse MLA support across decode, MTP, and DSpark, plus expanded AMD ROCm and NVFP4 support.
- –Qwen3.8 support on AMD ROCm broadens accelerator choice, while tiered KV-cache offloading and E/P/D disaggregation improve large-scale serving flexibility.
- –CUDA 13.0 becomes the default, but bitsandbytes migration and removed deprecated APIs mean production upgrades require compatibility testing.
DISCOVERED
8d ago
2026-08-29
PUBLISHED
8d ago
2026-08-29
RELEVANCE
AUTHOR
null_founder