AI labs appear to be optimizing their large language models (LLMs) specifically for a popular informal benchmark involving generating an SVG of a pelican riding a bicycle, according to a study published on July 18, 2026. Dylan Castillo conducted an experiment generating 1,008 SVGs across seven leading LLMs and analyzed the results with Claude Fable 5, revealing patterns consistent with targeted performance improvements.
The benchmark was originally created by Simon Willison, who has tested every major LLM release with the pelican-on-a-bicycle prompt. Castillo’s experiment involved scoring the SVG outputs with an LLM judge to assess quality and comparing the results across models. The analysis suggested that some AI labs may be 'pelicanmaxxing'—a term for optimizing models to excel on this specific test, potentially at the expense of broader capabilities.
This phenomenon raises questions about the reliability of informal benchmarks in evaluating AI progress, especially when billions of dollars are invested in model development. The pelican-on-a-bicycle prompt has become a well-known test within the AI community, often highlighted in Hacker News discussions. Castillo’s findings contribute to ongoing debates about whether labs prioritize benchmark performance over general model robustness.
All code and data from Castillo’s experiment are publicly available on GitHub, enabling further scrutiny and replication. The study was last updated on July 22, 2026, providing a timely insight into how AI labs approach evaluation metrics in competitive model development.