Liquid AI launches LFM2.5-VL-3B-DSpark drafter
Liquid AI released an experimental 279.5M-parameter speculative-decoding draft model for LFM2.5-VL-3B, accelerating token generation without changing outputs. Vendor benchmarks report up to 3.13× faster decoding on M5 Max and 2.66× on H100.
This is a smart inference optimization rather than a standalone model breakthrough, with especially strong implications for local multimodal applications. The main caveat is that results depend on a tightly coupled target model, runtime, and workload.
- –Adds only 8.9% more parameters while proposing blocks of tokens for the 3B target model to verify.
- –Supports SGLang, MLX-VLM, and llama.cpp across NVIDIA GPUs and Apple silicon.
- –End-to-end gains reach up to 2.62× on M5 Max and 2.27× on H100 because image encoding and prefill remain unchanged.
- –Faster local vision inference could improve screen understanding, OCR, visual agents, and interactive edge applications.
- –The model is experimental and target-specific, while reported benchmarks come from Liquid AI’s own testing.
DISCOVERED
1h ago
2026-09-26
PUBLISHED
2h ago
2026-09-26
RELEVANCE
AUTHOR
AiquestAcademy