Harness-IF exposes coding agents' default bias
Harness-IF is a research benchmark that tests whether coding agents actually follow rules across system prompts, project files, user instructions, tools, and skills. Across 12 model builds, aggregate compliance overstated performance on rules that opposed agents’ default behavior by an average of 5.81 points.
AGENTS.md guidance is only as reliable as the agent’s willingness to depart from its defaults, making rule-level, execution-based evaluation far more useful than simply checking task completion.
- –Evaluates 256 rules across 60 realistic multi-turn coding tasks and five instruction surfaces
- –Introduces Against-Prior Accuracy to separate genuine instruction following from behavior the model would have produced anyway
- –Every tested model performed worse on rules opposing its defaults, with gaps ranging from 3.6 to 7.4 points
- –Most failures were omissions of required actions, challenging the assumption that agents mainly violate prohibitions
- –The benchmark’s judge-heavy scoring means its broad patterns are more convincing than small leaderboard rank differences
DISCOVERED
1h ago
2026-08-13
PUBLISHED
1h ago
2026-08-13
RELEVANCE
AUTHOR
omarsar0