12 ms·
OpenAI o1 Results on ARC-AGI-Pub
- lossolo 2y agoIt seems like o1 is a lot worse than Claude on coding tasks https://livebench.ai https://livebench.ai
- devit 2y agoAm I missing something or this "ARC-AGI" thing is so ludicrously terrible that it seems to be completely irrelevant? It seems that the tasks consists of giving the model examples of a transformation of an input colored grid into an output colored grid, and then asking it to provide the output for a given input. The problem is of course that the transformation is not specified, so any answer is actually acceptable since one can always come up with a justification for it, and thus there is no reasonable way to evaluate the model (other than only accepting the arbitrary answer that the authors pulled out of who knows where). It's like those stupid tests that tell you "1 2 3 ..." and you are supposed to complete with 4, but obviously that's absurd since any continuation is valid given that e.g. you can find a polynomial that passes for any four numbers, and the test maker didn't provide any objective criteria to determine which algorithm among multiple candidates is to be preferred. Basically, something like this is about guessing how the test maker thinks, which is completely unrelated to the concept of AGI (i.e. the ability to provide correct answers to questions based on objectively verifiable criteria). And if instead of AGI one is just trying to evaluate how the model predicts how the average human thinks, then it makes no sense at all to evaluate language model performance by performance on predicting colored grid transformations. For instance, since normal LLMs are not trained on colored grids, it means that any model specifically trained on colored grid transformations as performed by humans of similar "intelligence" as the ARC-"AGI" test maker is going to outperform normal LLMs at ARC-"AGI", despite the fact that it is not really a better model in general.
- YeGoblynQueenne 2y agoNo no, that's not right. They're not asking for specific solutions. Any transformation of one grid to another will do.
- devit 2y agoWhat? They say: "ARC-AGI tasks are a series of three to five input and output tasks followed by a final task with only the input listed. Each task tests the utilization of a specific learned skill based on a minimal number of cognitive priors. Tasks are represented as JSON lists of integers. These JSON objects can also be represented visually as a grid of colors using an ARC-AGI task viewer. A successful submission is a pixel-perfect description (color and position) of the final task's output." As far as I can tell, they are asking to reproduce exactly the final task's output.
- YeGoblynQueenne 2y agoWhat they mean by "specific learned skill" is that each task illustrates the use of certain "core knowledge priors" that François Chollet has claimed are necessary to solve said tasks. You can find this claim in Chollet's white paper that introduced ARC, linked below: On the Measure of Intelligence https://arxiv.org/abs/1911.01547 https://arxiv.org/abs/1911.01547 "Core knowledge priors" are a concept from psychology and cognitive science as far as I can tell. To be clear, other than Chollet's claim that the "core knowledge priors" are necessary to solve ARC tasks, as far as I can tell, there is no other reason to assume so and every single system that has posted any above-0% results so far does not make any attempt to use that concept, so at the very least we can know that the tasks solved so far do not need any core knowledge priors to be solved. But, just to be perfectly clear: when results are posted, they are measured by simple comparison of the target output grids with the output grids generated by a system. Not by comparing the method used to solve a task. Also, if I may be critical: you can find this information all over the place online. It takes a bit of reading I suppose, but it's public information.
- deleted 2y ago[deleted]
- fancyfredbot 2y agoI found the level headed explanation of why log linear improvements in test score with increased compute aren't revolutionary the best part of this article. That's not to say the rest wasn't good too! One of the best articles on o1 I've read.
- Alifatisk 2y agoThis is a great marketing for Anthropic
- alphabetting 2y agoThis is best AGI benchmark out there in my opinion. Surprising results that underscore how good Sonnet is.
- zone411 2y agoDisagree. My opinion is that solving ARC-AGI won't get us any closer to AGI and it's mostly a distraction.
- alphabetting 2y agoHow so? I think if a team is fine-tuning specifically to beat ARC that could be true but when you look at Sonnet and o1 getting 20%, I think a standalone frontier model beating it would mean we are close or already at AGI.
- authorfly 2y agoThe creation and iteration of ARC has been designed in part to avoid this. Francis talks in his "mid-career" work (2015-2019) about priors for general intelligence and avoiding allowing them. While he admits ARC allows for some priors, it was at the time his best reasonable human effort in 2019 to put together and extremely prior-less training set, as he explained on podcasts around that time (e.g. Lex Fridman). The point of this is that humans, with our priors, are able to reliably get the majority of the puzzles correct, and with time, we can even correct mistakes or recognise mistakes in submissions without feedback (I am expanding on his point a little here based on conference conversations so don't take this as his position or at least his position today). 100 different humans will even get very different items correct/incorrect. The problem with AI getting 21% correct is that, if it always gets the same 21% correct, it means for 79% of prior-less problems, it has no hope as an intelligent system. Humans on the other hand, a group of 10000 could obviously get 99% or 100% correct despite none of them having priors for all of them in all liklihood given humans don't tend to get them all right (and well - because Francis created 100% of them!). The goal of ARC as I understood it in 2019, is not to create a single model that gets a majority correct, to show AGI, it has to be an intelligent system, which can handle prior or priorless situations, as good as a group of humans, on diverse and unseen test sets, ideally without any finetuning or training specifically on this task, at all. From 2019 (I read his paper when it came out believe it or not!), he held a secret set that he alone has that I believe is still unpublished, and at the time the low number of items (hundreds) was designed to prevent effective finetuning(then 'training') but nowadays few shot training shows that it is clearly possible to do on-the-spot training, which is why in talks Francis gave, I remember him positing that any advanced in short term learning via examples should be ignored e.g. each example should be zero shot, which I believe is how most benchmarks are currently done. The puzzles are all "different in different ways" besides the common element of dynamic grids and providing multiple grids as input. It's also key to know Francis was quite avant-garde in 2019: his work was ofcourse respected, but he became more prominent recently. He took a very bullish/optimistic position on AI advances at the time (no doubt based on keras and seeing transformers trained using it), but he has been proven right.
- GaggiX 2y agoIt really shows how far ahead Anthropic is/was when they released Claude 3.5 Sonnet. That being said, the ARC-agi test is mostly a visual test that would be much easier to beat when these models will truly be multimodal (not just appending a separate vision encoder after training) in my opinion. I wonder what the graph will look like in a year from now, the models have improved a lot in the last one.
- threeseed 2y ago> I wonder what the graph will look like in a year from now, the models have improved a lot in the last one. Potentially not great. If you look at the AIME accuracy graph on the OpenAI page [1] you will notice that the x-axis is logarithmic. Which is a problem because (a) compute in general has never scaled that well and (b) semiconductor fabrication will inevitably get harder as we approach smaller sizes. So it looks like unless there is some ground-breaking research in the pipeline the current transformer architecture will likely start to stall out. [1] https://openai.com/index/learning-to-reason-with-llms/ https://openai.com/index/learning-to-reason-with-llms/
- GaggiX 2y agoI'm very optimistic about it because native multimodal LLMs have hardly been explored. Also in general, I have yet to see these models plateau, Claude 3.5 Sonnet is a day and night different compared to previous models.
- accountnum 2y agoIt's not a problem, because the point at which we are in the logarithmic curve is the only thing that matters. No one in their right mind ever expected anything linear, because that would imply that creating a perfect oracle is possible. More compute hasn't been the driving factor of the last developments, the driving factor has been distillation and synthetic data. Since we've seen massive success with that, I really struggle to understand why people continue to doomsay the transformer. I hear these same arguments year after year and people never learn.
- fsndz 2y agoAs expected, I've always believed that with the right data allowing the LLM to be trained to imitate reasoning, it's possible to improve its performance. However, this is still pattern matching, and I suspect that this approach may not be very effective for creating true generalization. As a result, once o1 becomes generally available, we will likely notice the persistent hallucinations and faulty reasoning, especially when the problem is sufficiently new or complex, beyond the "reasoning programs" or "reasoning patterns" the model learned during the reinforcement learning phase. https://www.lycee.ai/blog/openai-o1-release-agi-reasoning https://www.lycee.ai/blog/openai-o1-release-agi-reasoning
- wslh 2y agoSo basically it's a kind of overfitting with pattern matching features? This doesn't undermine the power of LLMs but it is great to study their limitations.
- skepticATX 2y agoMy feeling is that this is one reason they decided to hide the reasoning tokens.
- fsndz 2y agoyes indeed
- poopiokaka 2y ago“As expected I’m right”
- fsndz 2y agoshouldn't I expect to be right when I have a thesis ? doesn't mean I can't see when I am wrong.
- meowface 2y agoTakeaway: >o1-preview is about on par with Anthropic's Claude 3.5 Sonnet in terms of accuracy but takes about 10X longer to achieve similar results to Sonnet. Scores: >GPT-4o: 9% >o1-preview: 21% >Claude 3.5 Sonnet: 21% >MindsAI: 46% (current highest score)
- deleted 2y ago[deleted]
- krackers 2y agoThere were rumors that 3.5 Sonnet heavily used synthetic data for training, in the same way that OpenAI plans to use o1 to train Orion. Maybe this confirm it?
- GaggiX 2y agoThe takeaway is also that o1-preview is a major improvement compare to GPT-4o. Anthropic is just ahead.
- Alifatisk 2y agoHow the hell is Anthropic this far ahead? I am yet impressed
- meowface 2y agoTrue. I've updated my post to include some of the scores.
- disgruntledphd2 2y agoIt's a little embarrassing for OpenAI though?
- dr_dshiv 2y agoI think they’d expect as much from Dario. He designed GPT3…
- attentive 2y ago
- w4 2y ago> o1's performance increase did come with a time cost. It took 70 hours on the 400 public tasks compared to only 30 minutes for GPT-4o and Claude 3.5 Sonnet. Sheesh. We're going to need more compute.
- Davidzheng 2y agoIntelligence is something that gets monotone easier as compute increases and trivial at the large compute limit (for instance can brute force simulate a human at large enough compute). So increasing compute is the most sure way to ensure success at reaching above human level intelligence (agi)
- trehalose 2y agoHow does one "brute force simulate a human"? If compute is the limiting factor, then isn't it currently possible to brute force simulate a human, just extremely slowly?
- soared 2y agoSomething something monkey at a typewriter writing Shakespeare
- rrrix1 2y agoGet out of my head!
- Davidzheng 2y agoThis is a more water tight proof of the same fact (so we don't have to argue about physics)
- azan_ 2y agoIt's not a proof at all.
- tomohelix 2y ago
- Terretta 2y agoTL;DR (direct quote): “In summary, o1 represents a paradigm shift from "memorize the answers" to "memorize the reasoning" but is not a departure from the broader paradigm of fitting a curve to a distribution in order to boost performance by making everything in-distribution.” “We still need new ideas for AGI.”
- sashank_1509 2y agoThis sounds very fair, but I think fundamentally humans memorize reasoning a lot more than you’d expect. A spark of inspiration is not memorized reasoning, but not many people can claim to enjoy that capability.
- benreesman 2y agoThe test you really want is the apples-to-apples comparison between GPT-4o faced with the same CoT and other context annealing that presumably, uh, Q* sorry Strawberry now feeds it (on your dime). This would of course require seeing the tokens you are paying for instead of being threatened with bans for asking about them. Compared to the difficulty in assembling the data and compute and other resources needed to train something like GPT-4-1106 (which are staggering), training an auxiliary model with a relatively straightforward, differentiable, well-behaved loss on a task like "which CoT framing is better according to human click proxy" is just not at that same scale.
- Stevvo 2y ago"Greenblatt" shown with 42% in the bar chart is GPT-4o with a strategy: https://substack.com/@ryangreenblatt/p-145731248 https://substack.com/@ryangreenblatt/p-145731248 So, how well might o1 do with Greenblatt's strategy?
- mikeknoop 2y agoI bet pretty well! Someone should try this. It's likely expensive but sampling could give you confidence to keep going. Ryan's approach costs about $10k to run the full 400 public eval set at current 4o prices -- which is the arbitrary limit we set for the public leaderboard.
- mrcwinn 2y agoHow is Anthropic accomplishing this despite (seemingly) arriving later?What advantage do they have?
- changoplatanero 2y agoOne theory I heard is that Dario was always interested in RL whereas Ilya was interested in other stuff until more recently. So Anthropic could have had an earlier start on some of this latest RL stuff.
- WiSaGaN 2y agoAnthropic currently does much less hype stuff comparing to openai. It's remarkable that openai was like this until the GPT-4 release, and completly changed since Sam Altman started touring countries.
- Satam 2y agoI think it's because OpenAI's leadership lacks good taste and talent. Realistically, they haven't shifted the needle with anything really interesting in 2 years now. They're using the inertia well but that's about it. Their model is not the best, the UI is not the best, and their pace of improvement is not great either.
- falcor84 2y agoI find the chatgpt-4o advanced mode to absolutely be "really interesting". And the video input they showed in the demos (and hope would same day release) could be a real game changer. One thing I would like to try, once that's out, is to put a computer with it amongst a group of students listening to a short lecture about something outside its training set and then check how the AI does on a comprehension quiz following the lecture - my feeling is that it would do significantly better than the average human student on most subjects.
- killthebuddha 2y agoIn my opinion this blog post is a little bit misleading about the difference between o1 and earlier models. When I first heard about ARC-AGI (a few months ago, I think) I took a few of the ARC tasks and spent a few hours testing all the most powerful models. I was kind of surprised by how completely the models fell on their faces, even with heavy-handed feedback and various prompting techniques. None of the models came close to solving even the easiest puzzles. So today I tried again with o1-preview, and the model solved (probably the easiest) puzzle without any kind of fancy prompting: https://chatgpt.com/share/66e4b209-8d98-8011-a0c7-b354a68fabca https://chatgpt.com/share/66e4b209-8d98-8011-a0c7-b354a68fab... Anyways, I'm not trying to make any grand claims about AGI in general, or about ARC-AGI as a benchmark, but I do think that o1 is a leap towards LLM-based solutions to ARC.
- riku_iki 2y agoInteresting part if you check CoT output, the way it solved: it said the pattern is to make number of filled cells even in each row with neat layout, which is interesting side effect, but not what task was about. It is also referring on some "assistant", looks like they have some mysterious component in addition to LLM, or another LLM.
- kobe_bryant 2y agoSo it gives you the wrong answer and then you keep telling it how to fix it until it does? What does fancy prompting look like then, just feeding it the solution piece by piece?
- killthebuddha 2y agoBasically yes, but there's a very wide range of how explicit the feedback could be. Here's an example where I tell gpt-4 exactly what the rule is and it still fails: https://chatgpt.com/share/66e514d3-ca0c-8011-8d1e-43234391a031 https://chatgpt.com/share/66e514d3-ca0c-8011-8d1e-43234391a0... and an example using gpt-4o: https://chatgpt.com/share/66e515da-a848-8011-987f-71dab56446f0 https://chatgpt.com/share/66e515da-a848-8011-987f-71dab56446... I'd share similar examples using claude-3.5-sonnet but I can't figure out how to do it from the claud.ai ui. To be clear, my point is not at all that o1 is so incredibly smart. IMO the ARC-AGI puzzles show very clearly how dumb even the most advanced models are. My point is just that o1 does seem to be noticeably better at solving these problems than previous models.
- a_wild_dandan 2y agoThis tests vision, not intelligence. A reasoning test dependent on noisy information is borderline useless.
- falcor84 2y agoWhat's noisy about it? The input matrix is discrete and converting it into any sort of structured input is trivial.
- ec109685 2y agoWhy is this considered such a great AGI test? It seems possible to extensively train a model on the algorithms used to solve these cases, and some cases feel beyond what a human could straightforwardly figure out.
- riku_iki 2y agoI think huge advantage is that they keep eval tests private, so corps can't finetune them to model and claim breakthrough, which possibly happened with many other benchmarks.
- visarga 2y agoThere is a hidden test set with new puzzle types not seen in the open part. It's designed so that humans do well and AI models have a hard time.
- YeGoblynQueenne 2y ago"Designed" is not right. What gives "AI models" (i.e. deep neural nets) a hard time is that there are very few examples in the public training and evaluation set: each task has three examples. So basically it's not a test of intelligence but a test of sample efficiency. Besides which, it is unfair because it excludes an entire category of systems, not to mention a dominant one. If F. Chollet really believes ARC is a test of intelligence, then why not provide enough examples for deep nets or some other big data approach to be trained effectively? The answer is: because a big data approach would then easily beat the test. But if the test can be beaten without intelligence, just with data, then it's not a test of intelligence. My guess for a long time has been that ARC will fall just like the Winograd Schema challenge (WSC) [1] fell: someone will do the work to generate enough (tens of thousands) examples of ARC-like tasks, then train a deep neural net and go to town. That's what happened with the WSC. A large dataset of Winograd schema sentences was crowd-sourced and a big BERT-era Transformer got around 90% accuracy on the WSC [2]. Bye bye WSC, and any wishful thinking about Winograd schemas requiring human intuition and other undefined stuff. Or, ARC might go the way of the Bongard Problems [3]: the original 100 problems by Bongard still stand unsolved, but the machine learning community has effectively sidestepped them. Someone made a generator of Bongard-like problems [4], and while this was not enough to solve the original problems, everyone simply switched to training CNNs and reporting results on the new dataset [5]. We basically have no idea how to create a test for intelligence that computers cannot beat by brute force or big data approaches so we have no effective way to test computers for (artificial) intelligence. The only thing we know humans can do that computers can't is identify undecidable problems (like Barber Paradoxes i.e. statements of the form "this sentence is false", as in Gödel's second incompleteness theorem). Unfortunately we already know there is no computer that can ever do that, and even if we observe say ChatGPT returning the right answer we can be sure it has only memorised, not calculated it, so we're a bit stuck. ARC won't get us unstuck in any way shape or form and so it's just a distraction. _____________________ [1] https://en.wikipedia.org/wiki/Winograd_schema_challenge https://en.wikipedia.org/wiki/Winograd_schema_challenge [2] WinoGrande: An Adversarial Winograd Schema Challenge at Scale https://arxiv.org/abs/1907.10641 https://arxiv.org/abs/1907.10641 Although note the results are interpreted to mean LLMs are more or less memorising answers, which is right of course. [3] Index of Bongard Problems https://www.foundalis.com/res/bps/bpidx.htm https://www.foundalis.com/res/bps/bpidx.htm [4] Comparing machines and humans on a visual categorization test https://www.pnas.org/doi/abs/10.1073/pnas.1109168108 https://www.pnas.org/doi/abs/10.1073/pnas.1109168108 [5] 25 years of CNNs: Can we compare to human abstraction capabilities? https://arxiv.org/abs/1607.08366 https://arxiv.org/abs/1607.08366
- bulbosaur123 2y agoOk, I have a practical question. How do I use this o1 thing to view codebase for my game app and then simply add new features based on my prompts? Is it possible rn? How?
- perching_aix 2y agoIs it possible for me, a human, to undertake these benchmarks?
- terhechte 2y agoThere's examples on the homepage, and there's a link to the Kaggle notebook in the article. https://arcprize.org https://arcprize.org