YOU ARE VIEWING ONE ITEM FROM THE AICRIER FEED

vLLM 0.28.0 Optimizes Open-Weight Inference

AICrier tracks AI developer news across Product Hunt, GitHub, Hacker News, YouTube, X, arXiv, and more. This page keeps the article you opened front and center while giving you a path into the live feed.

// WHAT AICRIER DOES

7+

TRACKED FEEDS

24/7

SCRAPED FEED

Short summaries, external links, screenshots, relevance scoring, tags, and featured picks for AI builders.

vLLM 0.28.0 Optimizes Open-Weight Inference
OPEN LINK ↗
// 8d agoINFRASTRUCTURE

vLLM 0.28.0 Optimizes Open-Weight Inference

vLLM 0.28.0, released August 26, delivers a major Kimi K3 performance push alongside deeper DeepSeek V4 and Qwen3.8 support.

// ANALYSIS

This release makes model architecture inseparable from serving economics: open-weight winners increasingly depend on the stack that turns exotic designs into affordable tokens. The gains remain hardware- and workload-dependent, so teams should benchmark their exact concurrency, context, quantization, and accelerator mix.

  • Kimi K3 gains fused FlashKDA kernels, decode-context parallelism, adaptive speculative decoding, and expert sharding that can save about 17 GiB per GPU.
  • vLLM reports up to 3.14× higher Kimi K3 throughput with DSpark on 16 NVIDIA GB300 GPUs. [Kimi K3 benchmarks](https://vllm-project.github.io/2026/07/27/k3.html)
  • DeepSeek V4 receives end-to-end sparse MLA support across decode, MTP, and DSpark, plus expanded AMD ROCm and NVFP4 support.
  • Qwen3.8 support on AMD ROCm broadens accelerator choice, while tiered KV-cache offloading and E/P/D disaggregation improve large-scale serving flexibility.
  • CUDA 13.0 becomes the default, but bitsandbytes migration and removed deprecated APIs mean production upgrades require compatibility testing.
// TAGS
vllminferenceopen-weightsopen-sourcemoequantizationgpu

DISCOVERED

8d ago

2026-08-29

PUBLISHED

8d ago

2026-08-29

RELEVANCE

9/ 10

AUTHOR

null_founder