3 ms·
Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.
by drob518 1mo ago
Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.
- gpt5 1mo agoWe should just consider the pelican bench as saturated and mostly meaningless.
- simondotau 1mo agoBut the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable, or a pink tennis ball with strawberry seeds and a reset button) and you can see that improvements are not narrowly tailored.
- crimsoneer 1mo agoSomeone tested this, and it doesn't look to be saturated. https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world. https://simonwillison.net/2026/Jul/16/kimi-k3/ https://simonwillison.net/2026/Jul/16/kimi-k3/
- kaoD 29d ago> Someone tested this, and it doesn't look to be saturated. They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>". Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
- simondotau 1mo agoThey're still not yet at the point where pelicanmaxxing is the best way to win this benchmark. Earlier models sucked because their SVG skills sucked. Newer models are likely better because more/better SVG models are being added to their training data.
- marci 1mo agoAfter looking at freely available SVGs of pelicans and bicycles, I have a hard time imagining what they could be using to game this.