Benchmark reveals AI agents ignore long policy documents
HANDBOOK.md is a benchmark consisting of 65 agentic tasks designed to evaluate whether language models can successfully complete complex work while adhering strictly to long policy documents. The study demonstrates that current frontier models perform poorly, with the best configurations achieving only a 36.2% pass rate, indicating that simply placing policies in the context window is insufficient for reliable governance.
Long-context windows might allow models to "read" large policy documents, but they completely fail at ensuring agents actually follow those rules when distracted by complex, multi-step tool use. The failure patterns are highly consistent: agents frequently let a plausible user request override a strict standing policy, and often perform required compliance checks but then inexplicably act against the results. These findings suggest that for enterprise deployments, critical policies cannot just be prompted; they must be enforced deterministically through tool guardrails.
DISCOVERED
2h ago
2026-07-29
PUBLISHED
5h ago
2026-07-29
RELEVANCE
AUTHOR
spIrr