5 ms·
I'm excited for the big jump in ARC-AGI scores from recent models, but no one should think for a second this is some leap in "general intelligence". I joke to
by mNovak 8mo ago
I'm excited for the big jump in ARC-AGI scores from recent models, but no one should think for a second this is some leap in "general intelligence".
I joke to myself that the G in ARC-AGI is "graphical". I think what's held back models on ARC-AGI is their terrible spatial reasoning, and I'm guessing that's what the recent models have cracked.
Looking forward to ARC-AGI 3, which focuses on trial and error and exploring a set of constraints via games.
- throw310822 8mo agoThe average ARC AGI 2 score for a single human is around 60%. "100% of tasks have been solved by at least 2 humans (many by more) in under 2 attempts. The average test-taker score was 60%." https://arcprize.org/arc-agi/2/ https://arcprize.org/arc-agi/2/
- modeless 8mo agoWorth keeping in mind that in this case the test takers were random members of the general public. The score of e.g. people with bachelor's degrees in science and engineering would be significantly higher.
- throw310822 8mo agoRandom members of the public = average human beings. I thought those were already classified as General Intelligences.
- thesmtsolver2 8mo agoAverage human beings with average human problems.
- imiric 8mo agoWhat is the point of comparing performance of these tools to humans? Machines have been able to accomplish specific tasks better than humans since the industrial revolution. Yet we don't ascribe intelligence to a calculator. None of these benchmarks prove these tools are intelligent, let alone generally intelligent. The hubris and grift are exhausting.
- throw310822 8mo ago> Machines have been able to accomplish specific tasks... Indeed, and the specific task machines are accomplishing now is intelligence. Not yet "better than human" (and certainly not better than every human) but getting closer.
- imiric 8mo ago> Indeed, and the specific task machines are accomplishing now is intelligence. How so? This sentence, like most of this field, is making baseless claims that are more aspirational than true. Maybe it would help if we could first agree on a definition of "intelligence", yet we don't have a reliable way of measuring that in living beings either. If the people building and hyping this technology had any sense of modesty, they would present it as what it actually is: a large pattern matching and generation machine. This doesn't mean that this can't be very useful, perhaps generally so, but it's a huge stretch and an insult to living beings to call this intelligence. But there's a great deal of money to be made on this idea we've been chasing for decades now, so here we are.
- warkdarrior 8mo ago> Maybe it would help if we could first agree on a definition of "intelligence", yet we don't have a reliable way of measuring that in living beings either. How about this specific definition of intelligence? Solve any task provided as text or images. AGI would be to achieve that faster than an average human.
- throw310822 8mo agoI still can't understand why they should be faster. Humans have general intelligence, afaik. It doesn't matter if it's fast or slow. A machine able to do what the average human can do (intelligence-wise) but 100 times slower still has general intelligence. Since it's artificial, it's AGI.
- guelo 8mo agoWhat's the point of denying or downplaying that we are seeing amazing and accelerating advancements in areas that many of us thought were impossible?
- colordrops 8mo agoWouldn't you deal with spatial reasoning by giving it access to a tool that structures the space in a way it can understand or just is a sub-model that can do spatial reasoning? These "general" models would serve as the frontal cortex while other models do specialized work. What is missing?
- causal 8mo agoThat's a bit like saying just give blind people cameras so they can see.
- pixl97 8mo agoI mean, no not really. These models can see, you're giving them eyes to connect to that part of their brain.
- amelius 8mo agoThey should train more on sports commentary, perhaps that could give spatial reasoning a boost.
- causal 8mo agoAgreed. I love the elegance of ARC, but it always felt like a gotcha to give spatial reasoning challenges to token generators- and the fact that the token generators are somehow beating it anyway really says something.
- deleted 8mo ago[deleted]