Long-Transduction exposes long-horizon agent failures
NVIDIA researchers introduce Long-Transduction, a controlled diagnostic for testing whether models can repeatedly read, mutate, and output state-dependent results across long contexts. Across open-weight models, performance fell 62.8% as context expanded from 4K to 128K, with additional losses from input-format variation and task complexity.
Long context is not the same as reliable long-horizon execution, and this benchmark makes that gap painfully measurable.
- –Tests arithmetic, UUID sorting, variable lookup, and table transformation across 1,440 controlled documents
- –Stable IDs and checkpointed subtasks substantially reduce failures caused by position tracking and output drift
- –At 128K, even the strongest tested model completed only 17.1% of documents without an error
- –Models often preserve task structure while copying the wrong information, suggesting retrieval alignment—not basic reasoning—is the core failure
- –Its synthetic design limits real-world conclusions, but exposes exactly the bookkeeping weaknesses agent workflows must engineer around
DISCOVERED
1h ago
2026-10-02
PUBLISHED
1h ago
2026-10-02
RELEVANCE
AUTHOR
Discover AI