Simon Willison: Are AI Labs Pelicanmaxxing? Research Across 48 Animal-Vehicle Combos Says No
Main idea
Dylan Castillo conducted systematic research asking: are AI labs benchmaxxing their models specifically for the famous informal pelican-on-bicycle test? Simon Willison highlights the experiment and its conclusion: no evidence of pelicanmaxxing was found.
Context
Simon Willison regularly tests every new LLM release with one prompt: generate an SVG of a pelican riding a bicycle. This benchmark became so well-known that the question arose whether AI labs might start deliberately optimizing for it (benchmaxxing). Castillo generated 1,000+ SVGs across 7 models and 48 prompts (8 animals x 6 vehicles).
Why it matters
The result is reassuring, but the existence of the question highlights the real risk of benchmark gaming in AI. For researchers and evaluation designers, it is a reminder that even informal benchmarks can become optimization targets — a phenomenon with direct impact on the credibility of AI evaluations.
Details / arguments
- Models tested: 7 frontier models
- Combinations: 48 (8 animals x 6 vehicles), each 3 times = 1,000+ SVGs
- Conclusion: no statistically significant pelicanmaxxing effects
- Pelicans are not drawn better than other animals individually
- Pelican+bicycle combination does not outperform prediction from individual scores