4 ms·
We should just consider the pelican bench as saturated and mostly meaningless.
by gpt5 1mo ago
We should just consider the pelican bench as saturated and mostly meaningless.
- simondotau 1mo agoBut the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable, or a pink tennis ball with strawberry seeds and a reset button) and you can see that improvements are not narrowly tailored.
- crimsoneer 1mo agoSomeone tested this, and it doesn't look to be saturated. https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world. https://simonwillison.net/2026/Jul/16/kimi-k3/ https://simonwillison.net/2026/Jul/16/kimi-k3/
- kaoD 1mo ago> Someone tested this, and it doesn't look to be saturated. They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>". Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.