Qwen3.8-27B NVFP4 Delivers 2.5× Speed
A two-GX10 document test found Qwen3.8-27B in NVFP4 matched BF16 quality while generating output 2.5× faster. The result strengthens NVFP4’s case as the default for local Blackwell inference when workloads tolerate quantization.
The compelling takeaway is not simply that four-bit inference is faster; it is that a workload-specific quality check erased BF16’s primary advantage. NVFP4 now looks like a practical capacity-and-throughput strategy for local AI developers, though one document benchmark cannot establish universal parity.
- –Unsloth lists NVFP4 at about 2.5× faster than BF16, with estimated memory falling from 56GB to 17–19GB for Qwen3.8-27B. [Unsloth](https://unsloth.ai/models/qwen3.8-27b)
- –The result is especially relevant for GX10, DGX Spark, RTX 50-series, and other Blackwell-class systems with native FP4 acceleration.
- –BF16 remains the safer baseline for sensitive evaluations, long-tail prompts, fine-tuning, and workloads where small numerical differences can affect correctness.
- –Developers should reproduce the comparison with fixed prompts, decoding settings, context lengths, kernels, and speculative-decoding configurations before generalizing the result.
- –Qwen3.8-27B became available on August 14, making this an early real-world inference benchmark rather than a model-launch announcement. [Qwen](https://github.com/QwenLM/Qwen3.8)
DISCOVERED
46d ago
2026-08-25
PUBLISHED
46d ago
2026-08-24
RELEVANCE
AUTHOR
Klausbjarner