4 ms·
What surprises me about this is how poorly general-purpose LLMs do. The best one is OpenAI o1-preview at 18%. This is significantly worse than the purpose-built
by celeritascelery 2y ago
What surprises me about this is how poorly general-purpose LLMs do. The best one is OpenAI o1-preview at 18%. This is significantly worse than the purpose-built models like ARChitects (which scored 53.5). This model used TTT to train on the ARC-AGI task specification (amoung other things). It seems that even if someone creates a model that can "solve" ARC, it still is not indicative of AGI since it is not "general" anymore, it is just specialized to this particular task. Similar to how chess engines are not AGI, despite being superhuman at chess. It will be much more convincing when general models not trained specifically for ARC can still score well on it.
They do mention that some of the tasks here are susceptible to brute force and they plan to address that in ARC-AGI-2.
> nearly half (49%) of the private evaluation set was solved by at least one team during the original 2020 Kaggle competition all of which were using some variant of brute-force program search. This suggests a large fraction of ARC-AGI-1 tasks are susceptible to this kind of method and does not carry much useful signal towards general intelligence.
- fchollet 2y agoIt is correct that the first model that will beat ARC-AGI will only be able to handle ARC-AGI tasks. However, the idea is that the architecture of that model should be able to be repurposed to arbitrary problems. That is what makes ARC-AGI a good compass towards AGI (unlike chess). For instance, current top models use TTT, which is a completely general-purpose technique that provides the most significant boost to DL model's generalization power in recent memory. The other category of approach that is working well is program synthesis -- if pushed to the extent that it could solve ARC-AGI, the same system could be redeployed to solve arbitrary programming tasks, as well as tasks isomorphic to programming (such as theorem proving).
- scoobertdoobert 2y agoFrançois, have you coded and tested a solution yourself that you think will work best?
- optimalsolver 2y agoHey, he's the visionary. You come up with the nuts and bolts.
- homarp 2y agois keras nuts and bolts enough?
- ipunchghosts 2y agoKeres is a good abstraction model but poorly implemented.
- ipunchghosts 2y ago"However, the idea is that the architecture of that model should be able to be repurposed to arbitrary problems" From a mathematical perspective, this doesn't sound right. All NNs are universal apprxomators and in theory can all learn the same thing to equal ability. It's more about the learning algorithm than the architecture IMO.
- mrandish 2y ago> It seems that even if someone creates a model that can "solve" ARC, it still is not indicative of AGI since it is not "general" anymore I recently explained why I like ARC to a non-technical friend this way: "When an AI solves ARC it won't be proof of AGI. It's the opposite. As long as ARC remains unsolved I'm confident we're not even close to AGI." For the sake of being provocative, I'd even argue that ARC remaining unsolved is a sign we're not yet making meaningful progress in the right direction. AGI is the top of Everest. ARC is base camp.
- iwsk 2y agoin other words, solving ARC is necessary but not sufficient for AGI
- mrandish 2y agoYes! That's the exact phrase I would have used with someone on HN. But that doesn't describe my non-technical friend. :-)
- YeGoblynQueenne 2y agoWhy is it necessary? Could a spider solve ARC-AGI, or could a pigeon, or a cat? And if an animal doesn't need to solve ARC-AGI to be intelligent, then why does an AGI?
- thomasahle 2y ago> What surprises me about this is how poorly general-purpose LLMs do. The best one is OpenAI o1-preview at 18%. o1-preview doesn't even have image input, so I wonder how they used it. Also, Ryan Greenblatts solution basically does "best of 4000" iirc. Presumably o1-preview was single shot.
- celeritascelery 2y agoNone of the models use images, they all operate and a json format the describes the input squares.