Horizon flags hidden benchmark task failures
Horizon’s new Data Quality Index reports that at least 29 of 239 tasks across widely used AI benchmark datasets are fundamentally flawed. The finding highlights how broken verifiers, underspecified success conditions, and brittle environments can distort agent evaluations.
Benchmark scores are only as trustworthy as the tasks behind them, and this audit suggests the field has prioritized scale over validity.
- –Defective tasks can measure verifier quirks or infrastructure reliability instead of model capability.
- –Common failure modes include instruction–verifier mismatches, impossible environment states, and overly rigid grading criteria.
- –Developers should expect task-level quality scores, reproducible audit evidence, versioned environments, and independent reruns alongside aggregate pass rates.
- –Curated suites such as Harbor-Index point toward a better tradeoff: fewer tasks, stronger audits, and clearer evidence that failures are genuine.
DISCOVERED
53m ago
2026-09-22
PUBLISHED
1h ago
2026-09-22
RELEVANCE
AUTHOR
horizon_bench