3 ms·
The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
by arctic-true 1mo ago
The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
- _diyar 1mo agoI strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
- CamperBob2 1mo agoAt this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
- jaggederest 1mo agoI feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
- aesthesia 1mo agoScoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
- aesthesia 1mo agoSee the scoring docs: https://docs.arcprize.org/methodology https://docs.arcprize.org/methodology
- _superposition_ 1mo agoReally though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.