6 ms·
Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark
- program_whiz 1mo agoIs this AGI? I don't think I can score 100% on ARC AGI.
- Insanity 1mo agoDepends on which definition they’ll use today.
- chris_st 1mo agoYou'll find that those goalposts are very movable.
- mdp2021 1mo ago[dead]
- tiahura 1mo agoCould KITT? Cmdr. Data?
- m3kw9 1mo agoagi with a context of 250k-1mill tokens?
- deleted 1mo ago[deleted]
- inerte 1mo agoIt is until ARC-AGI-4. Maybe around 73 we will stop.
- adastra22 1mo agoYes, we’ve had AGI for years now.
- program_whiz 1mo agoIts interesting because I didn't think it was, but then reading the NVIDIA approach, this kind of loop plus generating a program to explain things. Maybe that is AGI? I don't know, but it seems like an additional layer that maybe is a fundamental shift in capabilities (kind of like reinforcement learning and COT was).
- adastra22 1mo agoThe capability for a man-made machine machine to solve problems drawn from arbitrary problem domains without domain specific pre-training: Artificial. General. Intelligence. AGI. It is what the term of art means.
- daemonologist 1mo agoYou could probably score 100% on ARC 3 if you were motivated enough. I find some of the current problems to be kind of like Chess - mechanically simple, and ~solvable, but it's difficult to force myself to think at length about a monotonous and meaningless problem. The machines do have an advantage on the "energy" front; they've become almost psychotically persistent (and don't get tired after too many prompts). Anyway yes I think we've had AGI for a while now, even if the GI doesn't quite match up with what we expect from a human.
- edgarvaldes 1mo ago>A 100% score means AI agents can beat every game as efficiently as humans. (0) Yey, AGI is finally solved. (0) https://arcprize.org/arc-agi/3 https://arcprize.org/arc-agi/3
- andriy_koval 1mo ago> Is this AGI? I don't think I can score 100% on ARC AGI. 100% is some "RHAE" metric: its performance of median human first time seeing those problem.
- oblio 1mo agoWhat's AGI, at the end of the day? Equivalence to the average human?
- magicalhippo 1mo agoThe blog post: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/ https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-... Using Claude Opus 5, but it can use others: AVO is also designed to operate across frontier models. While our full public-set result used Claude Opus 5, we additionally paired AVO with GPT-5.6 Sol on a challenging subset of games. In these limited experiments, Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions in matched-level comparisons. These preliminary results suggest complementary operating profiles across models, and we leave a broader systematic comparison to future work
- AndrewKemendo 1mo agoNow we’re talking The next year is going to be wild folks
- tiahura 1mo agoAVO: Agentic Variation Operators for Autonomous Evolutionary Search https://arxiv.org/html/2603.24517v1 https://arxiv.org/html/2603.24517v1
- woeirua 1mo agoI thought ARC-AGI-3 was explicitly a test of raw model performance excluding the harness? Adding the harness back in doesn't tell us anything new. We've known that agents are capable of long horizon reasoning with sufficient harnesses. GPT-4(?) was capable of beating Pokemon 18 months ago but models only became capable of beating it without a harness in the last six months...?
- dist-epoch 1mo agoThe intent in forbidding harnesses was to prevent an ARC-AGI specific harness, which for example presented the game interface in a more agent-friendly way. What NVIDIA has here is a generic "evolution" harness, which can be used for any problem. I think it would be fair game to allow OpenClaw, Hermes, Codex, Grok Bot, this NVIDIA thing, to compete, as long as they don't have ARC-AGI specific skills, toolset.
- altcognito 1mo agoI'm assuming you're referring to a harness that includes memory -- I generally think of the harness as anything beyond executing the generation loop, but I'm not an expert. True as that may be, it may be better to optimize models for some amount of memory versus forcing some token count based on a reasoning level, right?
- throwaway314155 1mo agoI don't think GPT-4 was ever used for beating pokemon with or without a harness. Successful attempts include Gemini 2.5 and Opus 4.7 both using relatively advanced custom harnesses that give access to game memory, notes systems, and one-off hacks to get around parts of the game the model gets stuck on. More recently, Fable 5 beat FireRed with a _very_ minimal harness (screenshots and button inputs). That's the only example I know of but that is a very sophisticated and very expensive model compared to GPT-4. Most of this doesn't discredit your overall point, though.
- dezgeg 1mo agoGPT-o3 was the first OpenAI one to beat Red. The harness used by GPT Plays Pokemon is the most featureful one of the main competitors (GPT, Claude, Gemini), IIRC. Community maintained spreadsheet of the runs: https://docs.google.com/spreadsheets/d/e/2PACX-1vQDvsy5Dt_-Pg2PGe6LXRM8lokpUn4y6DQ4ShQLQPCGw5AOCPDG42pGnFfMOoqFU7eb7mPfHoGIB_c1/pubhtml#gid=1821408135 https://docs.google.com/spreadsheets/d/e/2PACX-1vQDvsy5Dt_-P...
- embedding-shape 1mo agoNone of the tweets, nor the press release, seems to mention how long time it actually took E2E to complete the evaluation, but they do mention it took "12% fewer actions" compared to just Opus 5 without AVO. Feels a bit suspicious they don't break down the timing involved, looking at the diagram from the press release, it gives the impression there is a lot of machinery here, and given they claim fewer actions, each action must be more carefully considered, doesn't it? Curious to read more about it though, seems the paper for it is here: https://arxiv.org/pdf/2603.24517 https://arxiv.org/pdf/2603.24517, I'm not sure I understand if it's better than just Codex with a /goal, as they talk about "can discover performance-critical micro-architectural optimizations" but leave Codex alone for a day or two and you'll get the same results without doing "additional autonomous adaptation" at all.
- subzel0 1mo agoThe 100% score was achieved on the 25 public set, not on the semi-private or private sets.
- xnx 1mo agoVerified high score is just 40%: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- moralestapia 1mo agoThat's a different thing. No AVO in there.
- throwaway2027 1mo agoI wonder if these benchmarks swap words, meaning and more because you might as well be benchmaxxing for specific words. I notice a lot of recurring just structural sentences coming back in smaller LLM models where they're fit for a specific task which is fine because most of the work we do is repetitive and there are patterns to learn but they should be word agnostic which I wonder if LLM can really fix.
- ru552 1mo agoThoughts on Nvidia releasing AVO or even open sourcing it? They've been very open with their models.
- deleted 1mo ago[deleted]
- angoragoats 1mo agoCan we please prioritize links to the papers, github repos, press releases, or blog articles for these types of posts? I don't use Twitter and I don't think anyone else should either.
- mellosouls 1mo agoUnderlying article should be the link: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/ https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-... There you will find the extremely important qualifier it's the public set, not the private set (with the risk of overfitting, ie the results not repeating when submitted to be run in competition), and the detail that this is essentially a harness added to Opus 5, not Nvidia's own models. Obviously still impressive, you would think.
- alok-g 1mo agoOnce ARC-AGI-3 is solved (including on the private set), would we be convinced that we have achieved AGI? If not, is the benchmark just incorrectly named? (I personally think so.) PS: I follow Wikipedia's definition for AGI (https://en.wikipedia.org/wiki/Artificial_general_intelligence https://en.wikipedia.org/wiki/Artificial_general_intelligenc...), which also talks about some tests. However, I distinguish it from Strong AI.
- kelseyfrog 1mo ago> would we be convinced that we have achieved AGI? No. AGI is impossible without a biological pineal gland. The pineal gland is the seat of consciousness and without one, any AI is merely a pattern matcher, not intelligent.
- alok-g 1mo agoYou seem saying: (1) Consciousness is mandatory for intelligence. (2) 'AGI' is no different than just 'intelligence'. (3) Consciousness resides in pineal gland. (4) Biology is mandatory for consciousness/AGI. I cannot claim these to be wrong, however, have no reason to believe in any. We do seem to agree though that the benchmark is incorrectly named.
- deleted 1mo ago[deleted]
- kelseyfrog 1mo ago[dead]
- wise_blood 1mo ago> Once ARC-AGI-3 is solved (including on the private set), would we be convinced that we have achieved AGI? no? once 3 is solved, we would come up with 4. then 5, 6... it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.
- antinucleon 1mo agoAVO’s paper author (ex-NVIDIAN) is here. This work was done half a year ago for GPU kernels, and the same approach has now been applied to ARC-AGI-3. I think people are still underestimating the evolution progress; e.g., recently we made a self-improving evolution harness that generated an entire inference stack and is better than SGLang/vLLM on various tasks: https://int21.ai/insights/addressing-the-inference-bottleneck/ https://int21.ai/insights/addressing-the-inference-bottlenec...