
Task-CoEvolve cuts harness eval costs 80%
Task-CoEvolve adaptively selects validation tasks where candidate agent harnesses disagree most, then estimates full-set performance from partial evaluations. Experiments on text classification and Terminal-Bench 2.1 matched full-search results while using substantially fewer evaluations.
This is a smart reframing of agent optimization: evaluation becomes an adaptive information-gathering problem instead of a brute-force tax.
- –Variance-weighted sampling focuses compute near the agent’s capability frontier
- –The method reportedly reaches comparable performance with roughly 20% of validation tasks per iteration
- –Terminal-Bench results show 67–80% lower search costs, though runtime savings are smaller because selected tasks can be longer
- –The approach could make iterative prompt, memory, retrieval, and tool-use optimization practical for more teams
- –Reproduction remains limited because the repository currently documents the method while code is still listed as coming soon
DISCOVERED
4h ago
2026-08-24
PUBLISHED
4h ago
2026-08-24
RELEVANCE
AUTHOR
Discover AI