A systematic study by Dylan Castillo tested 48 prompts across 7 AI models and found no evidence that labs trained their models specifically to draw pelicans on bicycles, debunking the 'pelicanmaxxing' hypothesis.

1 min read

Study Debunks 'Pelicanmaxxing' Hypothesis: No Evidence of Special Training

Study Debunks 'Pelicanmaxxing' Hypothesis

FAQ

What is the pelicanmaxxing hypothesis?

A hypothesis suggesting that AI labs may train their models specifically to draw pelicans on bicycles to improve performance on informal benchmarks.

How was the hypothesis tested?

48 prompts (8 animals × 6 vehicles) were tested 3 times each across 7 models, then evaluated using other models.

What is the main conclusion?

No evidence that any model draws pelicans on bicycles better than expected from its general capabilities.

Source: Simon Willison (LLM & tools)

AI-assisted content, human-reviewed.