6 ms·
Love seeing this benchmark become more iconic with each new model release. Still in disbelief at the GPT-5 variants' performance in comparison but its cool to s
by JJax7 11mo ago
Love seeing this benchmark become more iconic with each new model release. Still in disbelief at the GPT-5 variants' performance in comparison but its cool to see the new open source models get more ambitious with their attempts.
- aqme28 11mo agoOnly until they start incorporating this test into their training data.
- orbital-decay 11mo agoDataset contamination alone won't get them good-looking SVG pelicans on bicycles though, they'll have to either cheat this particular question specifically or train it to make vector illustrations in general. At which point it can be easily swapped for another problem that wasn't in the data.
- jug 11mo agoI like this one as an alternative, also requiring using a special representation to achieve a visual result: https://voxelbench.ai https://voxelbench.ai What's more, this doesn't benchmark a singular prompt.
- nwienert 11mo agothey can have some cheap workers make about 10 pelicans by hand in svg, fuzz them to generate thousands of variations and throw it in their training pool. don't need to 'get good at svgs' by any means.
- an0malous 11mo agoWhy is this a benchmark though? It doesn’t correlate with intelligence
- HighGoldstein 11mo agoWhat test would be better correlated with intelligence and why?
- ok_dad 11mo agoWhen the machines become depressed and anxious we'll know they've achieved true intelligence. This is only partly a joke.
- jiggawatts 11mo agoThis already happens! There have been many reports of CLI AI tools getting frustrated, giving up, and just deleting the whole codebase in anger.
- lukan 11mo agoThere are many reports of CLI AI tools displaying words that humans express when they are frustrated and about to give up. Just what they have been trained on. That does not mean they have emotions. And "deleting the whole codebase" sounds more interesting, but I assume is the same thing. "Frustrated" words lead to frustrated actions. Does not mean the LLM was frustrated. Just that in its training data those things happened so it copied them in that situation.
- jiggawatts 11mo agoThis is a fundamental philosophical issue with no clear resolution. The same argument could be made about people, animals, etc...
- lukan 11mo agoThe difference is, people and animals have a body, nerve system and in general those mushy things we think are responsible for emotions. Computers don't have any of that. And LLM's in particular neither. They were trained to simulate human text responses, that's all. How to get from there to emotions - where is the connection?
- K0balt 11mo agoI actually prefer ascii art diagrams as a benchmark for visual thinking, since it requires 2 stages, Like svg, and also can test imaginative repurposing of text elements.