Amazon ALoDLM releases adaptive diffusion models
Amazon’s ALoDLM family combines parallel diffusion-language generation with token-adaptive recurrent refinement, giving harder tokens more computation while easy tokens exit early. The released 1.7B and 8B Qwen3-derived models target research, reasoning, and code generation under a non-commercial license.
ALoDLM’s compelling idea is adaptive compute at the token level, addressing diffusion LMs’ quality gap without abandoning parallel decoding.
- –Reports 80.3 average performance across 11 benchmarks at 8B, exceeding the corresponding Qwen3 autoregressive baseline.
- –Reaches 93.25% on GSM8K at 612.4 tokens per second, roughly 2.7× faster than vLLM-served Qwen3-8B on one NVIDIA B200.
- –Includes research code and an optimized inference engine, but the speed claims depend on specialized hardware, kernels, and workload-specific settings.
- –The CC BY-NC 4.0 model license makes this valuable for experimentation while limiting straightforward commercial deployment.
- –If the approach generalizes, token-adaptive recurrence could make diffusion models more practical for latency-sensitive reasoning and coding workloads.
DISCOVERED
1h ago
2026-10-11
PUBLISHED
1h ago
2026-10-11
RELEVANCE
AUTHOR
AI Search