4 ms·
I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the r
by YmiYugy 2mo ago
I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted.
At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality.
We see a very janky pelican and declare the problem solved.
- twostorytower 2mo ago100%. Why waste the tokens to render Lord of the Rings when the pelican test still clearly benchmarks so well.
- jonas21 2mo agoI don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.
- Morromist 2mo agoI haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks. I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do. I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.
- enos_feedler 2mo agoAI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.
- Demiurge 2mo agoWhat you’re saying is that you’re not interested in benchmarks. But then you go a step further and state that this particular benchmark is entirely inconsequential. That’s like telling you that if you didn’t exist, nothing would change. Even if that were true, it would still be an insensitive and rude thing to say, wouldn’t it?
- Buttons840 2mo agoIt's okay to be rude to benchmarks though, they don't have feelings.
- Demiurge 2mo agoYeah, I agree, I didn’t mean to impose on this conversation between a man and a benchmark, my bad :p
- enos_feedler 2mo agoUsually when I say insensitive things there are more downvotes then upvotes. That isn't the case here. I might be rubbing against a truth somewhere here.
- w4yai 2mo ago> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle Please remember, we've started from there : https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.
- reaperducer 2mo agoSure, the task is not completed perfectly, but that's not the point. Isn't it? If the computer can't do it better than a human being, then what's the point? Being wrong at scale is not better than being right.
- matthewfcarlson 2mo agoMany humans would struggle with this even with very good tooling (ie not writing raw svg and using illustrator). I struggle to draw a bicycle accurately. But yes, I suspect it will be diminishing returns and I doubt it will ever be perfect due to the average nature of AI but I’d like to be wrong.
- skydhash 2mo ago> Many humans would struggle with this even with very good tooling But no ones hire random humans for things like this. You go and hire a vector artist and they will get your a very good pelican on a a bike. That's how you get things done when you can't do it.
- w4yai 2mo ago> You go and hire a vector artist Yeah, but then, you recruit the artist for $XXX - whereas you "recruit" your LLM for $0.XXX for the same task. Of course the quality difference is huge. But sometimes you don't need that level of quality. Also, finding a vector artist takes days of communication, payment settlement, revisions, etc. Not always the most practical solution.
- paulddraper 2mo agoPelicans don't ride bicycles. It's physically impossible. The problem is to draw it in the least disturbing way possible.
- Morromist 2mo agoTrue! But somehow Disney has been drawing ducks riding bikes in a way that seems to satisfy everyone since before my grandfather was born. https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and-louie-ride-a-bike/ https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...
- monk_grilla 2mo agoWow, that was a highly relevant and specific website to source here!
- Morromist 2mo agoI love these little websites with amazingly focused content. <3
- jodrellblank 2mo agoIt’s better than many of the AI offerings and the bike could steer and the ducks are sitting on saddles, but the three nephews can’t reach the bottom of the pedal stroke and by the looks of their feet on the far side of the bike they aren’t trying. That shouldn’t satisfy Donald and Daisy, leaving them with all the work.
- Demiurge 2mo agoAlso, one of the advantages of the pelican test is that you can evaluate it all at once. There’s no reason the two-dimensional depiction can’t be made more challenging. Yes, at some point the pelicans might approach the subjectivity of a fine art painting, but we haven’t even seen a depiction that’s competent by the standards of a high school art class. That’s not to say the elementary school–level SVGs aren’t amazing - rather, I agree with your point.
- BobbyJo 2mo agoI think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry? It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.
- Demiurge 2mo agoCould you share any examples that come close to those limits? I haven’t seen any that don’t have obvious flaws in proportions, layering, composition, color palette, visual clutter, or stylistic consistency.
- teiferer 2mo ago> touching the limit of what one can reasonably expect If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that. And Elon is already propagating the age of abundance where money won't exist anymore, right before calling the interviewing journalist dishonest and deservedly losing public trust. Smh my head.
- adriand 2mo ago> If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that. What knowledge workers do you know that have excellent drawing skills? I worked in a design agency and for a couple of years, each week me and a few other people would attempt to sketch a member of our group: one person would be the model and sit still, and everyone else would draw her/him. Let me tell you, if producing a convincing portrait was a prerequisite for being a knowledge worker, there would be 99% fewer knowledge workers.
- mattmanser 2mo agoHumans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design. So it's a pretty much pointless test now.
- teiferer 2mo agoYet no LLM can actually do it. It's quite surprising actually.
- rvz 2mo agoWhat does this even test for? Can I use LLMs to directly generate machine code to replace my compiler? Or maybe I can use LLMs as a bare metal OS / scheduler to replace my machine's operating system and scheduler? It makes zero sense to test for that. Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention. To end this obsession with generating SVG pelicans, Quiver AI's [0] model is actually designed to generate SVGs from prompts and has done so for years. So there is no need to continue with this un-serious pseudoscientific "benchmark". [0] https://quiver.ai/ https://quiver.ai/
- teiferer 2mo agoIt tests for the "I" in "AI".
- rvz 2mo agoSo if we test models that only output text to directly generate waveforms of sound from text or binary code, or directly generating binary code to replace a compiler, does that mean it is "intelligent"? Does that even test for intelligence? This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because each of them cannot fly or climb up trees.
- 2mo ago
- dllu 2mo agoAgreed. It's far from solved. Modern LLMs still generate pelican bike SVGs with obvious errors: * some omitted the bottom of the diamond which connects from the pedals to the rear wheel * some added an extra connection from the pedals to the front wheel, making it impossible to steer * none could align the head tube with the fork * none added a correct offset to the fork * none could generate the chain properly in a way that attaches to the two sprockets correctly I mean just look at these: * Grok 4.5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/pelicanmaxxing/x-ai__grok-4.5/pelican-bicycle__s0.png https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p... * GPT 5.6 Terra: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/pelicanmaxxing/openai__gpt-5.6-terra/pelican-bicycle__s0.png https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p... * Sonnet 5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/pelicanmaxxing/anthropic__claude-sonnet-5/pelican-bicycle__s0.png https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p... from https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html
- pj_mukh 2mo agowhat's fable at, does anyone know?
- dllu 2mo agohttps://simonwillison.net/2026/Jun/9/claude-fable-5/ https://simonwillison.net/2026/Jun/9/claude-fable-5/ The bike is generally okay, apart from medium which derped hard. Max has correct diamond, correct head tube, and so on. Only nitpicks are that the front fork offset isn't there and the chain doesn't touch the rear sprocket correctly.
- meander_water 2mo agoCan someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
- recursive 2mo agoIf your school mascot is a pelican and you need to make a flyer for the upcoming bike safety event, then this precise thing becomes useful. Most things that are useful in the real world seem not useful out of context.
- amelius 2mo agoYeah, pelican on a bicycle tells you how well the model can extrapolate outside of existing data, rather than just interpolate between it.
- meander_water 2mo agoSure, but then you would just use an image generation or multimodal model to generate that image. I don't think you'd want a weird looking svg.
- recursive 2mo agoExactly. I would want a good looking svg. That's the evaluation.
- RugnirViking 2mo agoOn the most basic level, its asking the ai to generate valid svg code for a picture of a pelican riding a bicycle, as a way of checking its intelligence. Popularized (invented?) by simonw, its been used as part of his reviews of new models as they come out since oct 2025. It used to be a very difficult task for models, see [2,3,4] it cuts across several tasks that AI used to be very bad at, but now has improved quite a bit. Namely, spatial reasoning (because it has to manually place the points of the svg such that they make sense and form what it says it forms. This used to not work very well, with random shapes floating around that it would mark things like "eyebrows" but were nowhere near the "eyes", etc. It also tests the model's world knowledge (what do pelicans look like? sure they have wings, feet, beaks etc, but what shape are they? how to get proportions roughly right? this isn't a given from text data about the bird. This goes doubly for a bike, which is a quite complex shape that most humans fail to draw correctly[1] (many draw the frame or chain connecting in impossible ways that would not ever function mechanically) Before it was pelican on a bicycle there were people having it do horses/unicorns making the rounds - gpt4.0 or whatever would often make hideous abominations of legs and mouths [1] https://www.gianlucagimini.it/portfolio-item/velocipedia/ https://www.gianlucagimini.it/portfolio-item/velocipedia/ [2] https://static.simonwillison.net/static/2026/mistral-small-4.png https://static.simonwillison.net/static/2026/mistral-small-4... [3] https://static.simonwillison.net/static/2025/codex-hacking-mini.png https://static.simonwillison.net/static/2025/codex-hacking-m... [4] https://static.simonwillison.net/static/2025/gemini-2.5-flash-lite-preview-09-2025.png https://static.simonwillison.net/static/2025/gemini-2.5-flas...
- crawfordmarch 2mo ago[flagged]