Liquid AI LFM2.5 DSpark Hits 3.18x
Liquid AI released DSpark draft checkpoints for LFM2.5-1.2B-Instruct, 2.6B, and 8B-A1B, enabling lossless speculative decoding. Benchmarks show up to 3.18× faster generation on an H100 and 2.87× on Apple M4 Max.
This is a practical inference win for edge and GPU deployments, though the headline speedup depends heavily on workload and token acceptance.
- –Draft models propose multiple tokens while the target verifies them in one forward pass
- –The 8B-A1B model reached 3.18× on MATH500 but only 1.29× on GSM8K
- –Output quality remains unchanged because rejected draft tokens are replaced by the target model
- –Official checkpoints cover three LFM2.5 sizes, with llama.cpp and SGLang deployment paths
- –Developers should benchmark their own prompts before assuming equivalent production gains
DISCOVERED
1d ago
2026-08-21
PUBLISHED
1d ago
2026-08-21
RELEVANCE
AUTHOR
Awesome_AI_News
