Back to section
Willison

Simon Willison: Are AI Labs Pelicanmaxxing? Research Across 48 Animal-Vehicle Combos Says No

Štvrtok 23. júla 2026 Source: simonwillison.net

Main idea

Dylan Castillo conducted systematic research asking: are AI labs benchmaxxing their models specifically for the famous informal pelican-on-bicycle test? Simon Willison highlights the experiment and its conclusion: no evidence of pelicanmaxxing was found.

Context

Simon Willison regularly tests every new LLM release with one prompt: generate an SVG of a pelican riding a bicycle. This benchmark became so well-known that the question arose whether AI labs might start deliberately optimizing for it (benchmaxxing). Castillo generated 1,000+ SVGs across 7 models and 48 prompts (8 animals x 6 vehicles).

Why it matters

The result is reassuring, but the existence of the question highlights the real risk of benchmark gaming in AI. For researchers and evaluation designers, it is a reminder that even informal benchmarks can become optimization targets — a phenomenon with direct impact on the credibility of AI evaluations.

Details / arguments

  • Models tested: 7 frontier models
  • Combinations: 48 (8 animals x 6 vehicles), each 3 times = 1,000+ SVGs
  • Conclusion: no statistically significant pelicanmaxxing effects
  • Pelicans are not drawn better than other animals individually
  • Pelican+bicycle combination does not outperform prediction from individual scores
Open original source simonwillison.net