HauhauCS pairs uncensoring with FastMTP
This community GGUF variant applies HauhauCS’s uncensored response profile to Qwen3.8-27B, while preserving multimodal capabilities, native MTP, and broad quantization options. An optional FastMTP sidecar targets substantially faster local inference.
FastMTP is the standout feature, but the uncensored profile makes this a specialized deployment choice rather than a general-purpose default.
- –The model spans roughly 10GB to 31GB quantizations, making local deployment practical across varied hardware.
- –HauhauCS reports up to 3.02x document throughput and 1.93x reasoning throughput versus MTP-disabled inference.
- –FastMTP requires a patched runtime, while embedded MTP works with current upstream llama.cpp.
- –The claimed zero refusals across 465 prompts may benefit research and experimentation, but removes safety behavior developers may expect.
- –Vision support requires the separate BF16 projector, adding deployment complexity for multimodal workloads.
DISCOVERED
1h ago
2026-09-26
PUBLISHED
2h ago
2026-09-26
RELEVANCE
AUTHOR
Oluwaphilemon1