CarWashBench flunks most frontier models

// 81d agoBENCHMARK RESULT

CarWashBench flunks most frontier models

CarWashBench v0.1 is a tiny public benchmark built around harder variants of the classic car-wash trick question to test whether LLMs can reason past surface cues. Across eight frontier models and five runs per question, only Gemini 3.1 Pro and GLM 5.0 showed meaningful performance while most models scored 0%.

// ANALYSIS

Tiny benchmarks can be noisy, but this one is a sharp gut-check for whether flagship "reasoning" models are actually reasoning or just pattern-matching.

–The benchmark is extremely small at just two questions, so the leaderboard is more signal flare than final verdict.
–Even so, near-total failure from multiple top-tier models is notable because the task targets everyday common-sense reasoning rather than specialized knowledge.
–Gemini 3.1 Pro standing well above the field suggests some models are better at escaping superficial heuristics on deceptively simple prompts.
–If the author expands the question set without losing the adversarial framing, this could become a useful lightweight reasoning stress test.

// TAGS

carwashbenchllmbenchmarkreasoningresearch

DISCOVERED

81d ago

2026-03-07

PUBLISHED

81d ago

2026-03-07

RELEVANCE

7/ 10

AUTHOR

Eyelbee

// KEEP READING

More AI developer news from the feed

EXPLORE FULL FEED

UPDATE1h ago

Cursor adds dedicated subagents for skills

Cursor now allows developers to execute tool-heavy or research-intensive agent skills within dedicated subagents. This architectural shift isolates noisy background tasks, keeping the main chat context clean and focused.

UPDATE1h ago

YouTube moves AI labels to video player

YouTube is moving its AI content disclosures from video descriptions to more prominent placements beneath the player and on Shorts overlays. Starting in May, the platform will use internal signals to automatically label photorealistic AI content that creators fail to disclose.

OPEN SOURCE5h ago

Taste Skill kills AI "frontend slop"

Taste-Skill is an open-source framework that provides portable "agent skills" to enforce high-end design principles in AI-generated code. By injecting specific design directives and "anti-slop" rules, it enables LLMs to produce editorial-grade UIs that bypass generic, boilerplate-heavy AI templates.