IBM BenchDrift exposes benchmark wording drift
IBM Research’s BenchDrift generates meaning-preserving variations of benchmark problems to measure how often models change answers because of phrasing. Tested across eight models and three benchmarks, the research finds that stronger models can be more vulnerable to wording changes.
BenchDrift makes a strong case that single-score leaderboards overstate model reliability.
- –Tests linguistic, referential, pragmatic, and structural variations while preserving the correct answer
- –Reveals both positive and negative drift, meaning rephrasing can fix failures or break correct answers
- –Shows phrasing sensitivity persists as models improve rather than disappearing
- –Gives developers a practical way to find prompt brittleness and hidden evaluation edge cases
- –Open-source code and data make robustness testing easier to reproduce
DISCOVERED
1h ago
2026-08-15
PUBLISHED
2h ago
2026-08-15
RELEVANCE
AUTHOR
omarsar0