Qwen3.8-27B EXL3 unlocks 200K local context
An experimental EXL3 quantization makes Qwen3.8-27B practical on consumer GPUs, including RTX 3090, 4090, and 5090 cards. Paired with DFlash2 and NVFP4 KV caching, the deployment kit targets roughly 200K–262K-token contexts within 24GB of VRAM.
This is a meaningful local-inference upgrade: smarter memory allocation and speculative decoding are making large open models increasingly viable on gaming hardware.
- –The 3.5bpw target weighs about 14.2GB, leaving room for long-context KV caching.
- –DFlash2 adds a 1.4GB draft model and can improve decoding speed, but reduces available context headroom.
- –NVFP4 KV caching is the key enabler for fitting extreme context windows into 24GB cards.
- –The stack includes an OpenAI-compatible server, ExLlamaV3 support, tool calling, and automatic model downloads.
DISCOVERED
1h ago
2026-09-01
PUBLISHED
1h ago
2026-09-01
RELEVANCE
AUTHOR
Oluwaphilemon1