4 ms·
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top
by jeswin 2mo ago
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
- farfatched 2mo agoIt's just a little fun.
- schafberg 2mo agobecuase most people don't care whether it's accurate, as long as it looks right and is funny...
- m00dy 2mo agoIt doesn't look right at all.
- SturgeonsLaw 2mo agoBut it does look funny
- zero0529 2mo agoWell it is just a bit of fun I think. However, I also think an AGI or an extremely capable model approaching AGI would be able to paint a pelican on a bicycle fairly easily. So in that way it is a good metric.
- rng-concern 2mo agoI agree it's fun, no argument there. However, it's no longer a good metric, as "drawing svg pelicans" is now showing up too much in the training data, so is not proof of generalization.
- pyaamb 2mo agoI prefer the Browser OS test
- amelius 2mo agoAt this point I want to see some human-drawn pelicans on bicycles. I suspect the LLMs aren't doing all that bad.
- setsewerd 2mo agoAs others commenters said, it's amusing. But also the person you're replying to is the guy who created the pelican test in the first place and I appreciate the whimsy he brings to the discussion.
- kanemcgrath 2mo agoBecause of all the svg rendering stuff, I added a draw_svg tool to my harness and it has been really nice to get a quick mock-up of ui changes. And conveniently, the new DeepSeek models are really good at knowing when to use it. So I do look at pelican rendering as a small metric of useful capability.
- bean469 2mo agoBecause seeing a pelican on a bike is always a good time. Look at it go
- Braden-dev 2mo agoI like to do something outlandish like a "cockatiel driving a UFO on it's way to austrailia"