TensorFold 1.0.3 Adds Qwen3.8-27B, DFlash2
TensorFold 1.0.3 adds native Qwen3.8-27B serving with DFlash2 speculative decoding, delivering up to 151 tokens per second on code workloads. The release also experiments with persistent model updates through Sliding Weights and adds native support for NVIDIA Ampere GPUs.
TensorFold is becoming a serious local-inference option, but the headline speedups depend heavily on using the drafter and compatible hardware.
- –Qwen3.8-27B replies match non-drafted output while substantially improving decode speed over stock mlx_lm.
- –Prompt processing still trails mlx_lm, and the Qwen engine serves one request at a time.
- –Sliding Weights provides intriguing model-level memory, but it rewrites checkpoint files and has no built-in undo.
- –Ampere support broadens TensorFold beyond Apple Silicon, though the release targets Nemotron rather than Qwen3.8-27B on NVIDIA GPUs.
- –Developers should benchmark time-to-first-token and prefill, not just decode throughput.
DISCOVERED
1h ago
2026-10-10
PUBLISHED
2h ago
2026-10-10
RELEVANCE
AUTHOR
Oluwaphilemon1

