SwiLA narrows softmax gap with constant-memory recall
Stanford researchers introduce SwiLA, a recurrent sequence layer that combines fixed-size linear-attention memory with dynamically routed linear regressors updated through online expectation-maximization. The paper reports stronger associative recall and competitive language-modeling results while keeping memory independent of sequence length.
SwiLA’s core idea is compelling: replace one limited linear memory with a mixture of specialized regressors, gaining context-dependent retrieval without restoring a growing KV cache.
- –SwiLA variants outperform DeltaNet and Mixture-of-Memories across six recall benchmarks, while hybrid SwiLA approaches softmax-attention performance.
- –Its effective mixture capacity grows as O(J^D), but stored state remains O(JD²), preserving sequence-length-independent memory.
- –The tradeoff is real: SwiLA is slower than GDN and has a larger fixed memory footprint, despite beating Transformer inference throughput.
- –Training currently relies on sequential Triton execution; efficient parallelization remains future work.
- –The official repository is MIT-licensed but still marked “work in progress,” so this is promising research rather than drop-in production infrastructure.
DISCOVERED
1h ago
2026-10-03
PUBLISHED
1h ago
2026-10-03
RELEVANCE
AUTHOR
Discover AI