4 ms·
No one’s talking about how good the final product is. Edit: someone else commented that as I was typing this, lol.
by demibabs 1mo ago
No one’s talking about how good the final product is.
Edit: someone else commented that as I was typing this, lol.
- blackhaz 1mo agoI wonder, do we need a new benchmark? There's quite a bit of feedback data floating around about pelicans on bicycles already.
- deleted 1mo ago[deleted]
- ziofill 1mo agoThat’s a fair question, but it seems that it’s not yet necessary. See here https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html https://simonwillison.net/2026/Jul/22/ https://simonwillison.net/2026/Jul/22/
- stymaar 1mo agoI don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”.
- Anon1096 1mo agoThat is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That's a general skill that is useful. So assuming they aren't specifically maximizing pelican bicycle svgs (and it doesn't look like they are) then incentives are aligned that the "benchmark" is measuring a general desirable capability.
- flexagoon 1mo agohttps://xkcd.com/810/ https://xkcd.com/810/
- simonw 1mo agoGemini have done exactly that. (I doubt it's because of my stupid benchmark, though!)