Qwen3.8 Coder390 Tames Runaway Reasoning
This open-weight Qwen3.8-27B variant uses additional SFT and RLOO post-training to reduce 94K-token truncations and empty answers while improving coding and reasoning benchmarks. Static FP8 reaches 178/198 on GPQA, 450/500 on MMLU, and 90/100 on LCB. [Model card](https://huggingface.co/nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2)
The meaningful breakthrough here is output discipline, not just another leaderboard bump: a reasoning model that knows when to stop is dramatically more useful in real coding workflows.
- –94K-token truncations fell from 4 to 1 on GPQA and from 13 to 3 on LCB.
- –Scores improved over the original Qwen3.8-27B, especially LCB, which rose from 83/100 to 90/100.
- –The model supports BF16, FP8, NVFP4, and GGUF variants, making local deployment practical across different hardware tiers.
- –DFlash2 and MTP speculative decoding reportedly reach about 180 and 91 tokens per second, respectively, versus roughly 46 without speculation.
- –The “Opus5.5” and “GPT6Astra” comparison is not independently established here; the published evaluation primarily compares against the original Qwen3.8-27B under a self-reported protocol.
DISCOVERED
2h ago
2026-10-11
PUBLISHED
2h ago
2026-10-11
RELEVANCE
AUTHOR
0x0SojalSec