4 ms·
There is still plenty of room for growth on the ARC-AGI benchmarks. ARC-AGI 2 is still <5% for o3-pro and ARC-AGI 1 is only at 59% for o3-pro-high: "ARC-AGI-1:
by croddin 1y ago
There is still plenty of room for growth on the ARC-AGI benchmarks. ARC-AGI 2 is still <5% for o3-pro and ARC-AGI 1 is only at 59% for o3-pro-high:
"ARC-AGI-1:
* Low: 44%, $1.64/task
* Medium: 57%, $3.18/task
* High: 59%, $4.16/task
ARC-AGI-2:
* All reasoning efforts: <5%, $4-7/task
Takeaways:
* o3-pro in line with o3 performance
* o3's new price sets the ARC-AGI-1 Frontier"
- https://x.com/arcprize/status/1932535378080395332 https://x.com/arcprize/status/1932535378080395332
- saberience 1y agoI’m not sure the arcagi are interesting benchmarks, for one they are image based and for two most people I show them too have issues understanding them, and in fact I had issues understanding them. Given the models don’t even see the versions we get to see it doesn’t surprise me they have issues we these. It’s not hard to make benchmarks that are so hard that humans and Lims can’t do.
- HDThoreaun 1y agoarc agi is the closest any widely used benchmark is coming to an IQ test, its straight logic/reasoning. Looking at the problem set its hard for me to choose a better benchmark for "when this is better than humans we have agi"
- saberience 1y agoThere are humans who cannot do arc agi though so how does an LLM not doing it mean that LLMs don’t have general intelligence? LLMs have obviously reached the point where they are smarter than almost every person alive, better at maths, physics, biology, English, foreign languages, etc. But because they can’t solve this honestly weird visual/spatial reasoning test they aren’t intelligent? That must mean most humans on this planet aren’t generally intelligent too.
- HDThoreaun 1y ago> LLMs have obviously reached the point where they are smarter than almost every person alive, better at maths, physics, biology, English, foreign languages, etc. I dont think memorizing stuff is the same as being smart. https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room > But because they can’t solve this honestly weird visual/spatial reasoning test they aren’t intelligent? Yes. Being intelligent is about recognizing patterns and thats what arc agi tests. It tests ability to learn. A lot of people are not very smart.
- ben_w 1y ago> I dont think memorizing stuff is the same as being smart. https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room I agree. The problem I have with the Chinese Room thought experiment is: just as the human who mechanically reading books to answer questions they don't understands does not themselves know Chinese, likewise no neuron in the human brain knows how the brain works. The intelligence, such as it is, is found in the process that generated the structure — of the translation books in the Chinese room, of the connectome in our brains, and of the weights in an LLM. What comes out of that process is an artefact of intelligence, and that artefact can translate Chinese or whatever. Because all current AI take a huge number of examples to learn anything, I think it's fair to say they're not particularly intelligent — but likewise, they can to an extent make up for being stupid by being stupid very very quickly. But: this definition of intelligence doesn't really fit "can solve novel puzzle", as there's a lot of room for getting good at that my memorising lot of things that puzzle-creators tend to do. And any mind (biological or synthetic) must learn patterns before getting started: the problem of induction* is that no finite number of examples is ever guaranteed to be sufficient to predict the next item in a sequence, there is always an infinite set of other possible solutions in general (though in reality bounded by 2^n, where n = the number of bits required to express the universe in any given state). I suspect, but cannot prove, that biological intelligence learns from fewer examples for a related reason, that our brains have been given a bias by evolution towards certain priors from which "common sense" answers tend to follow. And "common sense" is often wrong, c.f. Aristotelian physics (never mind Newtonian) instead of QM/GR. * https://en.wikipedia.org/wiki/Problem_of_induction https://en.wikipedia.org/wiki/Problem_of_induction
- nipah 1y ago"most people I show them too have issues understanding them, and in fact I had issues understanding them" ??? those benchmarks are so extremely simple they have basically 100% human approval rates, unless you are saying "I could not grasp it immediately but later I was able to after understanding the point" I think you and your friends should see a neurologist. And I'm not mocking you, I mean seriously, those are tasks extremely basic for any human brain and even some other mammals to do.
- clbrmbr 1y agoYou may be above average intelligence. Those challenges are like classic IQ tests and I bet have a significant distribution among humans.
- achierius 1y agoNo, they've done testing against samples from the general population.
- yorwba 1y agoThe ARC-AGI-2 paper https://arxiv.org/pdf/2505.11831#figure.4 https://arxiv.org/pdf/2505.11831#figure.4 uses a non-representative sample, success rate differs widely across participants and "final ARC-AGI-2 test pairs were solved, on average, by 75% of people who attempted them. The average test-taker solved 66% of tasks they attempted. 100% of ARC-AGI-2 tasks were solved by at least two people (many were solved by more) in two attempts or less." Certainly those non-representative humans are much better than current models, but they're also far from scoring 100%.
- cubefox 1y agoThe original ARC-AGI test was much easier than the recent v2.
- saberience 1y agolol 100% approval rates? No they don’t. Also mammals? What mammals could even understand we were giving it a test? Have you seen them or shown them to average people? I’m sure the people who write them understand them but if you show these problems to average people in the street they are completely clueless. This is a classic case of some phd ai guys making a benchmark and not really considering what average people are capable of. Look, these insanely capable ai systems can’t do these problems but the boys in the lab can do them, what a good benchmark.