NEEDLE benchmark exposes web search blind spots
Keenable’s new open-source NEEDLE benchmark refreshes live search tasks hourly or daily, making memorization ineffective. Its current leaderboard gives Keenable 75.7% of pooled best-possible performance, versus 53.1% for Google’s API.
The result is a sharp reminder that search infrastructure—not just model intelligence—sets an agent’s ceiling.
- –NEEDLE covers news, finance, scholarly retrieval, legal search, and rare long-tail queries.
- –Engines receive identical queries and are judged on returned titles, snippets, ranking, and answer recall.
- –The 75.7% figure represents share of an oracle assembled from all providers, not coverage of the entire internet.
- –Google remains strong on known-item lookups but falls behind on fresh, obscure, and agentic queries.
- –Because the benchmark is public, continuously refreshed, and open source, providers have less room to optimize for a fixed leaderboard.
DISCOVERED
1h ago
2026-08-27
PUBLISHED
1h ago
2026-08-27
RELEVANCE
AUTHOR
EXM7777