MiMo Distill 9B Adds GPU-Friendly Quants
Hikari07jp has added Q4_K_M and NVFP4 quantization options to its refusal-ablated MiMo 9B model, including a 5.63GB build aimed at smaller GPUs. NVFP4 W4A8 reportedly reaches 108 tokens per second on an RTX 5070 Ti. [Hugging Face](https://huggingface.co/Hikari07jp/MiMo-V2.6-Distill-Qwen-9B-Ablitrated)
This makes an unusual local model substantially easier to run, but the convenience comes with meaningful evaluation and safety caveats.
- –Q4_K_M brings the model into roughly 6GB territory, making 12GB GPUs much more practical targets.
- –NVFP4 W4A8 and W4A4 offer alternative memory and throughput trade-offs for supported hardware.
- –The base model’s ablation removes refusal behavior, but its own audit found most removed refusals became soft deflections rather than substantive answers.
- –Quantization can further affect capability, so developers should benchmark coding, tool use, and safety behavior on their own workloads before deployment.
- –This is best viewed as an enthusiast-friendly local inference experiment, not a production-ready aligned assistant.
DISCOVERED
1h ago
2026-09-26
PUBLISHED
1h ago
2026-09-26
RELEVANCE
AUTHOR
Oluwaphilemon1