6 ms·
Hey HN, Greg from ARC Prize Foundation here. Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measur
by gkamradt 2y ago
Hey HN, Greg from ARC Prize Foundation here.
Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measures a model’s ability to generalize on tasks it hasn’t seen before, and the ARC Prize 2025 competition to beat it.
In Dec ‘24, ARC-AGI-1 (2019) pinpointed the moment AI moved beyond pure memorization as seen by OpenAI's o3.
ARC-AGI-2 targets test-time reasoning.
My view is that good AI benchmarks don't just measure progress, they inspire it. Our mission is to guide research towards general systems.
Base LLMs (no reasoning) are currently scoring 0% on ARC-AGI-2. Specialized AI reasoning systems (like R1 or o3-mini) are <4%.
Every (100%) of ARC-AGI-2 tasks, however, have been solved by at least two humans, quickly and easily. We know this because we tested 400 people live.
Our belief is that once we can no longer come up with quantifiable problems that are "feasible for humans and hard for AI" then we effectively have AGI. ARC-AGI-2 proves that we do not have AGI.
Change log from ARC-AGI-2 to ARC-AGI-2:
* The two main evaluation sets (semi-private, private eval) have increased to 120 tasks
* Solving tasks requires more reasoning vs pure intuition
* Each task has been confirmed to have been solved by at least 2 people (many more) out of an average of 7 test taskers in 2 attempts or less
* Non-training task sets are now difficulty-calibrated
The 2025 Prize ($1M, open-source required) is designed to drive progress on this specific gap. Last year's competition (also launched on HN) had 1.5K teams participate and had 40+ research papers published.
The Kaggle competition goes live later this week and you can sign up here: https://arcprize.org/competition https://arcprize.org/competition
We're in an idea-constrained environment. The next AGI breakthrough might come from you, not a giant lab.
Happy to answer questions.
- artninja1988 2y agoWhat are you doing to prevent the test set being leaked? Will you still be offering API access to the semi private test set to the big model providers who presumably train on their API?
- gkamradt 2y agoWe have a few sets: 1. Public Train - 1,000 tasks that are public 2. Public Eval - 120 tasks that are public So for those two we don't have protections. 3. Semi Private Eval - 120 tasks that are exposed to 3rd parties. We sign data agreements where we can, but we understand this is exposed and not 100% secure. It's a risk we are open to in order to keep testing velocity. In theory it is very difficulty to secure this 100%. The cost to create a new semi-private test set is lower than the effort needed to secure it 100%. 4. Private Eval - Only on Kaggle, not exposed to any 3rd parties at all. Very few people have access to this. Our trust vectors are with Kaggle and the internal team only.
- zamadatix 2y agoWhat prevents everything in 4 from becoming a part of 3 the first time the test set is run on a proprietary model, do you require competitors like OpenAI provide models Kaggle can self host for the test?
- gkamradt 2y ago#4 (private test set) doesn't get used for any public model testing. It is only used on the Kaggle leaderboard where no internet access is allowed.
- zamadatix 2y agoSorry, I probably phrased the question poorly. My question is more along the lines of "when you already scored e.g. OpenAI's o3 on ARC AGI 2 how did you guarantee OpenAI can't just look at its server logs to see question set 4"?
- gkamradt 2y agoAh yes, two things 1. We had a no-data retention agreement with them. We were assured by the highest level of their company + security division that the box our test was run on would be wiped after testing 2. We only tested o3 against the semi-private set. We didn't test it with the private eval.
- zamadatix 2y agoMakes sense, particularly part 2 until "the final results" are needed. Thanks for taking the time to answer my question!
- YeGoblynQueenne 2y ago>> We were assured by the highest level of their company + security division that the box our test was run on would be wiped after testing Yuri Geller assured us he was bending the spoons with his mind. Somehow it was only when the Amazing Randi was present that Yuri Geller couldn't bend the spoons with his mind.
- gmkhf 2y agoI think a lot of people got discouraged, seeing how openai solved arc agi 1 by what seems like brute forcing and throwing money at it. Do you believe arc was solved in the "spirit" of the challenge? Also all the open sourced solutions seem super specific to solving arc. Is this really leading us to human level AI at open ended tasks?
- jmtulloss 2y agoWhy is this the same comment as https://news.ycombinator.com/item?id=43466406 https://news.ycombinator.com/item?id=43466406?
- synapsomorphy 2y agoThanks for your awesome work Greg! The success of o3 directly contradicts us being in an "idea-constrained environment", what makes you believe that?
- littlestymaar 2y agoWhat makes you think so? From ChatGPT 3.5 to o1, all LLMs progress came from investment in training: either by using much more data, or using higher quality data thanks to artificial data. o1 (and then o3) broke this paradigm by applying a novel idea (RL+search on CoT) and that's because of it that it was able to make progress on ARC-AGI. So IMO the success of o3 goes in favor of the argument of how we are in an idea-constrained environment.
- torginus 2y agoThis isn't a novel idea - some people tried the exact same thing the day GPT4 came out. And going back even further, there's Goal Oriented Action Planning - an old timey video game AI technique, that's basically searching through solution space to construct a plan: https://medium.com/@vedantchaudhari/goal-oriented-action-planning-34035ed40d0b https://medium.com/@vedantchaudhari/goal-oriented-action-pla... (besides the fact that almost all old timey AI is state space solution search)
- littlestymaar 2y agoWhat's new is to apply that to LLMs, that is. > This isn't a novel idea - some people tried the exact same thing the day GPT4 came out. What do you mean? Since GPT4's weights aren't available, you can't run RL on it by yourself. Only OpenAI can.
- jononor 2y agoNot Greg/team, so unrelated opinion. o3 solution for ARC v1 was incredibly expensive. Some good ideas are at least needed to take that cost down by a factor 100-10000x.
- vessenes 2y agoJust want to say I really love these new problems - feels like some general intelligence went into conceiving of and creating these puzzles: we just did a few over dinner as a family. You have my wheels turning on how to get computers better at these. Looking forward to see G the first computer tech that can get 30-50% on these!
- az226 2y agoDid any single individual solve all problems? How many such individuals were there?
- levocardia 2y agoI'm really pleased to see this! The original ARC-AGI-1 paper still informs how I think about "what is intelligence" today. I was thrilled to see AI models make real progress on that test precisely when we had the next big idea (reasoning). Here's to hoping round 2 falls with a similarly big breakthrough!
- tananaev 2y agoDid I read this right that only 2 humans out of 400 solved the problems?
- trott 2y agoThey started with N >= 120x3 tasks, and gave each task to 4-9 humans. Then they kept only those 120x3 tasks that at least 2 humans had solved.
- tananaev 2y agoThat's a very small sample size by task. I wonder if they give the whole data set to an average human, what the result would be. I tried some simple tasks and they are doable, but I couldn't figure out the hard ones.
- mapmeld 2y agoNo, they're saying that the problems have been reviewed / play-tested by ≥2 humans, so they are not considered unfair or too ambiguous to solve in two attempts (a critique of some Arc-AGI-1 puzzles that o3 missed). They have a lot of puzzles so they were divided among some number of testers, but I don't think every tester had to try every problem.
- doctorpangloss 2y agoWhy doesn’t every blogpost contain an example of a question you ask?
- Chathamization 2y ago> Our belief is that once we can no longer come up with quantifiable problems that are "feasible for humans and hard for AI" then we effectively have AGI. I don’t think that follows. Just because people fail to create ARC-AGI problems that are difficult for an AI to solve, doesn’t mean that said AI can just be plugged into a humanoid robot and it will now reliably cook dinner, order a pizza and drive to pick it up, take a bus to downtown to busk on the street and take the money back home, etc. ARC-AGI is an interesting benchmark, but it’s extremely presumptive to think that these types of tests are going to demonstrate AGI.
- Palmik 2y agoIn your example you already indicated two tasks that you think might be hard for AI but easy for humans. Who said that cooking dinner couldn't be part of ARC-AGI-<N>?
- Chathamization 2y agoThat’s precisely what I meant in my comment by “these types of tests.” People are eventually going to have some sort of standard for what they consider AGI. But that doesn’t mean the current benchmarks are useful for this task at all, and saying that the benchmarks could be completely different in the future only underscores this.
- pillefitz 2y agoThey are useful to reach Arc-N+1
- Chathamization 2y agoHow are any of these a useful path to asking an AI to cook dinner? We already know many tasks that most humans can do relatively easily, yet most people don’t expect AI to be able to do them for years to come (for instance, L5 self-driving). ARC-AGI appears to be going in the opposite direction - can these models pass tests that are difficult for the average person to pass. These benchmarks are interesting in that they show increasing capabilities of the models. But they seem to be far less useful at determining AGI than the simple benchmarks we’ve had all along (can these models do everyday tasks that a human can do?).
- Centigonal 2y agoThank you for including cost (or really any proxy for efficiency) as a dimension to this prize!
- az226 2y agoWhich puzzles had the lowest solve rate? I did the first 10 and felt all easy (mentally solve it in 10-20 seconds for easier ones and 30-60 seconds for harder ones), I’d like to try the most difficult ones.
- celdon25 2y agoWhy wasn’t the ICOM framework (D. Kelley) allowed to make a scoring submission after they claimed to have beaten the scores? Are you concerned that may appear to contradict your mission statement and alienate the AGI community?
- ustad 2y agoUsing AGI in the titles of your tests might not be accurate or appropriate. May I suggest NAI - Narrow AI?
- JFingleton 2y agoMy prediction: we'll be arguing about what AGI actually is... Forever.
- throwuxiytayq 2y agoOr depending on your outlook, for a couple of years, and then we will no longer be participating in these or any other cognitive exercises.