3 ms·
Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
by atemerev 1mo ago
Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
- maipen 1mo agoVery well said. It kinda describes how unrealistic these expectations are. Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result. Very very unrealistic and wasteful.
- wieiw1 1mo ago[dead]
- droidjj 1mo agoThe complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
- saaaaaam 1mo agoWhen AGI comes it will come as a pelican and gobble up all these troublesome little fishies who gripe and whine and moan about pelicans.
- balefulboy 1mo agoI'm always tired of seeing at the top of every new model release post on here. I say Simon should just keep it to Twitter.
- saaaaaam 1mo agoI have never seen a Pelican on X which used to be called twitter in about 1872. Keep up!
- saaaaaam 1mo agoGood grief. You’re no fun either. This whole thread is lots of no fun. Pelicans are fun.
- saaaaaam 1mo agoWell you’re just no fun are you?!
- simonw 1mo agoHah, this is a new one: first time there's been a complaint about the pelican before I've even posted one! (I don't have access yet.)
- persedes 29d agoShower thought: But how often per week do you run the pelican these days? And do you have it automated at this point or would the automation take out the meaning of the benchmark?
- simonw 29d agoMy automation is pretty simple. I use my https://llm.datasette.io https://llm.datasette.io tool where I have a template saved: llm "Generate an SVG of a pelican riding a bicycle" --save pelican When a model comes out I first make sure LLM can talk to it - usually by updating the relevant plugin, but if it's on OpenRouter I can use it directly with https://github.com/simonw/llm-openrouter https://github.com/simonw/llm-openrouter - sometimes I use this mechanism instead, for OpenAI-compliant API models: https://llm.datasette.io/en/stable/other-models.html#configure-an-openai-compatible-model https://llm.datasette.io/en/stable/other-models.html#configu... Then I run something like this: llm -m gpt-6-astra -m pelican Then I grab the most recent log export as markdown: llm logs -cu | pbcopy -c means most recent conversation, -u includes token usage I paste that into https://gist.github.com https://gist.github.com and then paste the resulting Gist URL into the URL tab on https://tools.simonwillison.net/markdown-svg-renderer https://tools.simonwillison.net/markdown-svg-renderer If the model supports multiple reasoning levels I run it once per level and put those in the same file. I really should automate this a bit more.