4 ms·
Context: ARC Prize 2024 just wrapped up yesterday. ARC Prize's goal is to be a north star towards AGI. The two major categories of this year's progress seem to
by mikeknoop 2y ago
Context: ARC Prize 2024 just wrapped up yesterday. ARC Prize's goal is to be a north star towards AGI. The two major categories of this year's progress seem to fall into "program synthesis" and "test-time fine tuning". Both of these techniques are adopted by DeepMind's impressive AlphaProof system [1]. And I'm personally excited to finally see actual code implementation of these ideas [2]!
We still have a long way to go for the grand prize -- we'll be back next year. Also got some new stuff in the works for 2025.
Watch for the official ARC Prize 2024 paper coming Dec 6. We're going to be overviewing all the new AI reasoning code and approaches open sourced via the competition [3].
[1] https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/ https://deepmind.google/discover/blog/ai-solves-imo-problems...
[2] https://github.com/ekinakyurek/marc https://github.com/ekinakyurek/marc
[3] https://x.com/arcprize https://x.com/arcprize
- aithrowawaycomm 2y agoI am a bit uncertain about the rules of the ARC-AGI contest, but would this program count? A good chunk of the logic of ARC is essentially hardcoded, including a Python function that checks whether or not the proposed solution makes sense. The point of the contest is to measure intelligence in general-purpose AI systems: it does not seem in the spirit of the contest that this AI would completely fail if the test was presented on a hexagonal grid.
- 0x1064 2y agoThe point in the contest is to measure an algorithms ability to solve ARC problems specifically, no one believes that it's general-purpose AI. They're highly contrived problems by design.
- aithrowawaycomm 2y agoMy point is that the contest really should be "can solve ARC problems without having anything about ARC problems in its pre-training data or hard-coded in the design of the program." Otherwise these claims from ARC-AGI are simply false: Solving ARC-AGI represents a material stepping stone toward AGI. At minimum, solving ARC-AGI would result in a new programming paradigm. It would allow anyone, even those without programming knowledge, to create programs simply by providing a few input-output examples of what they want. This would dramatically expand who is able to leverage software and automation. Programs could automatically refine themselves when exposed to new data, similar to how humans learn. If found, a solution to ARC-AGI would be more impactful than the discovery of the Transformer. The solution would open up a new branch of technology. This program does not represent a "new paradigm" because it requires a bunch of human programming work specifically tailored to the problem, and it cannot be generalized. If software like this wins the contest that really shows the contest has nothing whatsoever to do with AGI.
- antonvs 2y ago> Otherwise these claims from ARC-AGI are simply false: Currently, any claim about AGI other than "we're probably not anywhere close to strong AGI" is simply false. Of course a lot depends on one's definition of AGI. From another perspective, one could argue that ChatGPT 4 and similar models are already AGI.
- benchmarkist 2y agoThe contest is misnamed, solving ARC will not get us any closer to AGI.
- naasking 2y agoWhy?
- benchmarkist 2y agoBecause it's a set of puzzles on a 2D grid. We don't live on a 2D grid so it's already on the wrong track. A set of puzzles for a 3D sphere wouldn't get us any closer to AGI either but at least it would be a more realistic representation of the world and how a general purpose problem solver should approach reality. Even Minecraft would be a better test and lately people have started testing LLMs in virtual worlds which is a much better test case than ARC. Insofar as ARC is being used as a benchmark for code synthesis it might be somewhat successful but it doesn't seem like people are using code synthesis to solve the puzzles so it's not really clear how much success on ARC is going to advance the state of the art in AI and code synthesis according to a logical specification.
- deleted 2y ago[deleted]
- razodactyl 2y agoYou're on track in your arguments but don't underestimate how hard the puzzles in ARC actually are. It takes a considerable amount of depth in reasoning to see and reason about the patterns / problems / solutions. Try doing a few of them by hand to see what I mean. Simulated worlds are complex enough to hide their own flaws just like LLMs are complex enough to lead us to believe they can reason when most of the time they are pattern matching.
- Zondartul 2y agoARC problems are too hard for me. I'm no longer sure I'm generally intelligent.
- razodactyl 2y agoMajority of ARC can be gamed / hard-coded, no doubt about it. The real pressure is the private hold-out set and the variations that can be added to counter this aspect. A true AGI would be able to solve anything thrown at it which is where the authors are trying to lead AI engineering towards since LLMs have pretty much taken over. If it starts getting too easy, they just reconsider and add harder problems. It's like how we don't talk about the Turing Test anymore as it's no longer the best metric to determine real intelligence. The authors are signalling to the industry that new ideas are needed and the monetary aspect is to show how serious they are about it. It's good because as per above we have research being thrown at it which means we can iterate until we perhaps find another breakthrough.