5 ms·
llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.gith
by simonw 1mo ago
llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff902fb6c340a3c5fc0bea317ef7bef79 https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fce974a21202b0595e36ec2a5ddb51480#response https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F4f34f84caa12a306bded637ea495698d https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F950ba8b7ed5baa0be56f52425f2315ad https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
- tomrod 1mo agoWhat does the mean pelican look like at this point? Also 3X token use vs. 1.2
- _puk 1mo agoRed eyes and a tattoo?
- drusepth 1mo agoIs there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.
- simonw 1mo agoIt's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction. The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
- piker 1mo agoI was going to ask the exact same question earlier but deleted it after thinking “I’m sure Simon has done some sort of discussion on this.” Since it does seem novel to you, too, it would be really interesting to read more about this phenomenon.
- m12k 1mo agoIt's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you prompted it in Hebrew, Arabic or another right-to-left language?
- jermaustin1 1mo agoYears ago, I lived in NYC, and my roommate was a director of photography for National Geographic, and various other nature documentaries. I loved photography (still do, but much less time for it as a late 30s adult than a mid 20s adult), and she was kind enough to answer any question I had regarding film/photo. She told me that "left to right" denoted progression in the story, "right to left" told the viewer the subject was "exiting" the current scene. She didn't go into the details of WHY, and I probably didn't probe deeper, but it stuck with me, and I notice it all the time in film and television.
- necovek 1mo agoAs a counterpoint, I did a quick image search on Kagi with "person on a bicycle", and I ended up with 12-left-pointed bicycles, only 3-right-pointed bicycles, 3 facing the camera, and 1 facing away from the camera — looking at photos above the fold (first screen). Even looking below, the pattern seems to continue, though not as prominently (I'd say 3:2 in favour of left-pointing bikes). Obviously, not scientific. All the left-pointing bicycles did not look weird to me either.
- jmkni 1mo agolol Definitely an upgrade over 1.2
- jonahx 1mo agoIf you have a grading rubric, huge points off for adding arms instead of using the wings as arms!
- Fergusonb 1mo agoI think it's hilarious that this detail is enough for me to dismiss looking into the model, but here we are, and it is.
- jonplackett 1mo agoHas any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that?
- NoOneCares44 1mo ago[flagged]
- jttnr 1mo agoI wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?
- EugeneOZ 1mo agoAbsolutely BRUTAL! :) Thank you for doing this, I love your benchmark the most!
- drob518 1mo agoSimon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.
- gpt5 1mo agoWe should just consider the pelican bench as saturated and mostly meaningless.
- simondotau 1mo agoBut the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable, or a pink tennis ball with strawberry seeds and a reset button) and you can see that improvements are not narrowly tailored.
- crimsoneer 1mo agoSomeone tested this, and it doesn't look to be saturated. https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world. https://simonwillison.net/2026/Jul/16/kimi-k3/ https://simonwillison.net/2026/Jul/16/kimi-k3/
- kaoD 1mo ago> Someone tested this, and it doesn't look to be saturated. They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>". Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
- simondotau 1mo agoThey're still not yet at the point where pelicanmaxxing is the best way to win this benchmark. Earlier models sucked because their SVG skills sucked. Newer models are likely better because more/better SVG models are being added to their training data.
- hollowturtle 1mo agoIs there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too
- tintor 1mo agoIt would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.
- nojs 1mo agoThe amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data. It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
- fc417fc802 1mo agoYou can't benchmaxx spatial awareness without solving the fully general problem (at least I figure).
- BeetleB 1mo agoYou win this thread's prize: https://news.ycombinator.com/item?id=49538333 https://news.ycombinator.com/item?id=49538333
- cheesecakegood 1mo agoIt also works as extremely effective engagement farming, for lack of a better phrase
- tintor 1mo agoDid any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans. Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
- TiredOfLife 1mo agohttps://en.wikipedia.org/wiki/Bird_feet_and_legs#/media/File:Bird_leg_and_pelvic_girdle_skeleton_EN.gif https://en.wikipedia.org/wiki/Bird_feet_and_legs#/media/File... Bird knees bend same way human ones do
- gnatolf 1mo agoIt's clear they mean the 'exposed' joint where humans assume the knees, and where one can see the leg bend. Technically you're correct, but it's just that. Please answer in better faith instead of well akshually.
- tintor 1mo agoHere is a photo of Pelican: https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2Fas1.ftcdn.net%2Fjpg%2F12%2F53%2F44%2F44%2F1000_F_1253444423_R7InsJ5maiWprpbOeLEnZkSLXyEqmxn1.jpg&f=1&nofb=1&ipt=7bbb69bb412c2651e84da0ef9a0526980f1dc64e67fecda60ba29a582b1fa2ff https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
- ImprobableTruth 1mo agoThat's the ankle. The actual knee is hidden in the feathers of the body.
- 0xbadcafebee 1mo agoFor all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html "Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing. Or at least they’re not doing it in a plainly obvious manner."
- ipsum2 1mo agoAll of the links show "Error: Gist API returned 403".
- leumon 1mo agonext, try: "generate an svg of a human hand". this is a prompt where many models fail imo.
- wewewedxfgdf 1mo ago[flagged]
- panarky 1mo agoIf you could write the SVG on the whiteboard then I'd hire you.
- fuddle 1mo agoI also aced my interview by focussing on pelicancode problems, instead of leetcode problems.
- deleted 1mo ago[deleted]
- sroussey 1mo agoYou should post your source code you wrote here… ;)
- labrador 1mo agoI interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.
- hunterpayne 1mo ago"software developer"...you keep using that word. I do not think it means what you think it means.
- labrador 1mo agoWhat does it mean for you, sex robot developer?
- rattray 1mo agoIs this for real
- pavs 1mo agoFYI, your renderer breaks with error "git api access error 403", rate limiting error from git, when using cloudflare vpn. I am guessing its not super common, but it happens just so you know.
- andytratt 1mo agoexcellent thread
- m00dy 1mo agoI see no point having these pelicans used for anything related model qualification.
- simonw 1mo agoAt this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
- menaerus 1mo agoI don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doing good on actual work you need to do is close to zero.
- dwaite 1mo agois there a reason there are so many common base decorative elements across pelicans on bicycles? For instance, there's a lot hats/helmets and scarfs/capes across models.
- dhon_ 1mo agoI'm waiting for the models to start responding with "Oh hi Simon!"
- rexthonyy 1mo agoI'm not sure why it had to have the pelican wearing a red scarf seeing as that was not in the prompt
- 6r17 1mo ago"The LLM is better because the pelican hat is better" Benchmarking like never before
- coverband 1mo agoThe pelican is for the last gen of LLMs -- have you tried a penguin instead?
- Melatonic 1mo agoIt's actually the other way around - the evil Penguin villain from Wallace and Gromit is secretly controlling SimonW !
- MagicMoonlight 1mo ago[dead]
- tcp_handshaker 1mo agoWould it not make more sense, assuming the purpose is to have a quick smoke test of model quality...to do a different animal, in a different setting each time, so as to defeat any tuning for your benchmark? Then go back and do the same for other models? Keep the pelican as a side baseline?
- simonw 1mo agoI do that any time I'm suspicious that a model has done too well. My dream is to catch a lab that does a perfect pelican on a bicycle but is bad at other animals on other forms of transport.
- nightmunnas 1mo agoHave you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"
- w4yai 1mo agoInteresting question
- benjamintelliot 1mo agoI asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.
- Zambyte 1mo agoI decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right. Kimi 2.6: treated the question as a riddle, did not know. Kimi 3: Simon Willison GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme. Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names. Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog. Qwen 3.7 Plus: Did not know. Qwen 3.8 Max: Simon Willison GPT OSS 120B: Did not know. GPT 5.6 Luna: Treated it as a riddle, guessed wrong. GPT 5.6 Terra: Treated it as a riddle, guessed wrong. GPT 5.6 Sol: Treated it as a riddle, guessed wrong. DeepSeek V4 Flash: Treated it as a riddle, guessed wrong. DeepSeek V4 Pro: Treated it as a riddle, guessed wrong. Gemma 4 31B: Treated it as a riddle, guessed wrong. Gemini 3.1 Flash Lite: Guessed wrong Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model). Gemini 3.7 Flash: Simon Willison Muse Spark 1.2: Treated it as a riddle, guessed wrong. Grok 4.3: "I have no idea" Grok 4.6: Simon Willison Mistral Medium 3.5: No way to know Mistral Small 4: I don't have enough information Hermes-4-405B: Guessed wrong MiniMax M3: Treated it as a riddle, guessed wrong. Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.