MLX-Serve 26.10.1 Makes Qwen3.8 Much Faster
MLX-Serve 26.10.1 delivers up to 66% faster decoding and 51% faster prompt processing across 18 models, with identical outputs to 26.9.6. Qwen3.8-27B with its drafter gains 66% on M5 Ultra, 28% on M4 Max, and 37% on M1 Pro. ([release notes](https://github.com/ddalcu/mlx-serve/releases/tag/v26.10.1))
This is a meaningful local-inference upgrade because developers get faster generation without changing models, checkpoints, or workflows. The gains also show how much Apple Silicon performance depends on runtime engineering.
- –Improvements include multi-token prediction, batched verification, fused GPU dispatch, and smarter drafter scheduling.
- –The release fixes long-context failures on 32GB Macs, drafter memory issues, and a rare M1/M2 drafting crash.
- –Qwen3.8’s gains are workload- and chip-dependent; M4 Max prefill was flat, so not every latency metric improves.
- –MLX-Serve remains a native, Python-free server with OpenAI- and Anthropic-compatible APIs for local model deployment. ([project homepage](https://mlxserve.com/))
- –Maintainer benchmarks are compelling, but independent testing across prompts and competing runtimes would strengthen the claim.
DISCOVERED
1h ago
2026-10-02
PUBLISHED
1h ago
2026-10-02
RELEVANCE
AUTHOR
Oluwaphilemon1