A pelican-on-a-bicycle SVG is a meaningful benchmark of overall AI reasoning
What's this about?
People disagree about whether asking AI to draw a pelican riding a bicycle tests overall AI reasoning.
This task can test some useful skills, but it cannot test every kind of thinking.
What supporters say
- The AI must join two ideas and show the pelican truly riding the bicycle.
- It must understand where each thing goes, not just place both things near each other.
- SVG code must make a real picture that people can change later.
- This task can test image-making, clear prompt-following, and code-writing skills.
What critics say
- One fun drawing task cannot show how well AI reasons about everything.
- The task does not directly test logic, plans, math, or cause and effect.
- It also does not test whether AI handles new problems when rules change.
- A picture that looks right may still come from messy code or weak understanding.
The bottom line
A pelican riding a bicycle can test a few real AI skills.
But evidence does not show that this one task measures overall AI reasoning.
A request to create an SVG of a pelican riding a bicycle can test some real AI skills. But the evidence does not support treating that one whimsical task as a stand-in for overall AI reasoning.
The case for
The prompt is not trivial. To succeed, an AI must combine separate ideas—a pelican and a bicycle—and correctly show the relationship between them: the bird must be riding, rather than merely appearing beside, above or inside the bicycle. That makes it a reasonable test of compositional understanding and relational accuracy, two abilities studied in text-to-image evaluation (see Figure 2). 1
Requiring the answer in SVG format adds another layer. SVG is structured vector code, not simply a written description or a flat image. A good response must produce a valid, editable graphic while preserving the requested scene. That can reveal whether a model can translate an instruction into a visual design with meaningful object structure, geometry and layering. 2
This matters because a model that repeatedly creates usable SVGs of the requested relationship is demonstrating more than word repetition. It is showing selected skills in text-to-image alignment, visual composition and structured code generation. These are legitimate and useful areas of AI evaluation.
The task could become more informative with clearer scoring. Evaluators could separately judge whether the SVG renders properly, whether it contains the right objects and attributes, whether the pelican is actually riding the bicycle, and whether the underlying code is editable and well structured. Such a system would be stronger than a simple judgment that the image “looks right.”
The case against
The main problem is scale. One open-ended illustration cannot measure overall reasoning, which includes many abilities the task does not directly test: deduction, planning, mathematics, causal understanding, transfer to new situations and reliability under changing conditions. Broader AI benchmarks typically use varied task types, real-world problems with functional tests, or unfamiliar visual challenges meant to test abstraction and generalization (see Figure 3). 3
A successful pelican drawing may be consistent with some reasoning ability, but it does not show that a system can reason well across those other domains. There is no evidence that a high score on this prompt predicts performance on independent measures of reasoning. That is the central gap in the claim.
Scoring is another weakness. If a human simply decides whether the output appears acceptable, the result can mix together several different questions: Did the AI follow the prompt? Is the image aesthetically pleasing? Is the bicycle clear? Is the SVG code valid? A polished-looking image may still contain hidden structural or relational errors, while many visually different images could meet the prompt’s basic requirements. 4
The fixed prompt may also reward familiarity rather than general ability. An AI could have encountered similar animal-on-vehicle images, common SVG patterns or well-known prompting conventions during training. Without testing paraphrases, unfamiliar animal-and-vehicle pairings, explicit spatial constraints or edits to an existing SVG, evaluators cannot tell whether the model is applying a transferable skill or reproducing learned patterns. 5
More controlled versions would help. Repeated trials, varied prompts and comparison with unrelated reasoning tasks could show whether performance generalizes. But even under those conditions, the benchmark would support conclusions about particular multimodal and code-generation abilities—not every form of reasoning.
The bottom line
The evidence supports a qualified negative conclusion. A pelican-on-a-bicycle SVG is a meaningful targeted test of compositional understanding, spatial relationships and structured vector-code generation. It is not a meaningful standalone benchmark of overall AI reasoning.
Confidence in that distinction is high. Research supports the value of compositional visual and SVG tasks, while broader claims about reasoning require diverse tasks, systematic scoring and tests of generalization. The key uncertainty is not whether this prompt measures anything; it is whether performance on a carefully controlled set of similar tasks would meaningfully track broader, independently measured reasoning ability.
Figures & data
All contributions are reviewed for clarity, balance, and evidence. The strongest insights are elevated into the argument graph — with credit to you.
Help improve this analysis →