Environment Evolution hardens terminal-agent training
Tencent’s Hunyuan team introduces an off-policy curriculum that evolves verified terminal environments across increasingly difficult generations. Long-horizon RL with the evolved tasks improved Qwen3.6-27B and Qwen3.6-35B-A3B by 14.4 and 18.0 percentage points on Terminal-Bench 2.1.
The paper targets a real bottleneck in agent training: once models outgrow static tasks, useful RL signals disappear. Its strongest idea is treating environment difficulty as a controllable training variable rather than endlessly mining model-specific failures.
- –Evolves tasks along scenario novelty, skill rarity, and execution length.
- –Uses verifier-gated multi-agent loops to preserve solvability while increasing difficulty.
- –An Evolution-Lineage Scheduler prevents exposing policies to environments they cannot yet solve.
- –Reported peaks of 71.5% and 64.9% outperform both co-evolution and environment-ensemble baselines.
- –Reproduction remains demanding: the study uses 500 curated seed environments, 200 GRPO steps, and long-context terminal-agent infrastructure.
DISCOVERED
1h ago
2026-09-06
PUBLISHED
1h ago
2026-09-06
RELEVANCE
AUTHOR
Discover AI