Study finds no AI pelican benchmark gaming
Researcher Dylan Castillo tested 1,008 SVG images across seven frontier models to evaluate whether AI labs secretly over-optimize for Simon Willison's viral 'pelican riding a bicycle' benchmark. Automated judging and regression modeling revealed that pelicans and bicycles score in the lower half of generated subjects with no statistically significant performance boost, showing SVG rendering gains reflect general model capabilities rather than benchmark maxxing.
Informal community meme benchmarks are fun for casual evaluation, but assuming AI labs are secretly overfitting to every viral prompt gives capability teams way too much credit for niche benchmark gaming.
- –Neither pelicans nor bicycles scored systematically higher than other animal or vehicle categories, with both ranking near the lower end of overall judge scores.
- –A fixed-effects regression adjusting for prompt difficulty found no statistically significant performance boost for the pelican-on-a-bicycle prompt cell across any of the seven models tested.
- –While all pelican-on-a-bicycle images faced right, directional analysis showed this composition bias reflects general side-profile drawing mechanics for bicycles rather than a memorized scene artifact.
- –Quality variations across models reflect general multimodal code-generation capability rather than targeted fine-tuning on community benchmark prompts.
DISCOVERED
4h ago
2026-07-22
PUBLISHED
7h ago
2026-07-22
RELEVANCE
AUTHOR
dcastm
