GLM-5.3 EXL3 Fits Four DGX Sparks
VCRUZ305’s SAGE MixedK quantization compresses GLM-5.3 to 319 GB while retaining all 19,200 routed experts and its 1M-token context. The checkpoint targets full-model local inference across four NVIDIA DGX Sparks.
This is an impressive memory-efficiency result, but its practical frontier depends on runtime maturity and real-world quality validation.
- –Keeps the complete MoE topology instead of pruning or merging experts, preserving fidelity at an unusually low 3.38 bpw.
- –Fits weights plus a full 1M-token KV cache within four 128 GB DGX Sparks.
- –Reported serving reaches up to 69 tokens per second with the specialized TensorFold and DFlash2 stack.
- –Stock ExLlamaV3 cannot load the checkpoint, making deployment tooling the immediate bottleneck.
- –Quality evidence is promising but limited; broader coding, reasoning, and long-context evaluations are still needed.
DISCOVERED
1h ago
2026-10-08
PUBLISHED
1h ago
2026-10-08
RELEVANCE
AUTHOR
ViC305