6 ms·
GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt
by meetpateltech 3mo ago
GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8%
Sol is the first verified frontier model to ever beat an ARC-AGI-3 game
https://arcprize.org/results/openai-gpt-5-6 https://arcprize.org/results/openai-gpt-5-6
- simianwords 3mo agoVery interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?
- osti 3mo agoMythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
- vatsachak 3mo agoAre you joking? They spend billions of dollars training LLMs to get a 7.8% on arc agi 3 whereas DINO models are near sota in image classification, provide meaningful embeddings to the point where image segmentation is just PCA. The spend on DINO cannot be more than five million (correct me if I'm wrong) JEPA is just getting started
- esafak 3mo agoASI is going to be here by the time Lecun gets started.
- redactsureAI 3mo agoDINO is a transformer model?
- vatsachak 3mo agoJEPA can use a transformer, and DINO does so yes
- deleted 3mo ago[deleted]
- Tenoke 3mo agoHis main anti-LLM predictions have been consistently either wrong or misleading. There's many ways to skin a cat so you can probably do something with a JEPA approach as well, but I doubt he actually catches up to having agents on the level of where Anthropic/OpenAI will be at any point.
- onlyrealcuzzo 3mo agoHis main LLM predictions have almost nothing to do with Arc AGI... What exactly was he dead wrong about that is proven by any of this? GPT getting better has absolutely nothing to do with completely disproving anything LeCun has been saying. He never said LLMs couldn't get better. He never said they couldn't score 7.6% on Arc AGI 3. He's merely said they don't think, and you probably want something that actually thinks if you want a model that can be trained cheaply on a small amount of data and provide a ton of value. Spending $5B to train a model that scores better than an older model does not disprove any of that in any way.
- Tenoke 3mo ago>He's merely said they don't think He said years ago even 'GPT 5000' couldnt do things that they ended up doing fine a month later, let alone by 5000. His later predictions are just moving that goal post including towards them not being able to do more general, harder problems of which Arc AGI is a counter-example.
- onlyrealcuzzo 3mo ago> He said even 'GPT 5000' couldnt do things that they could do a month later, let alone by 5000. What things specifically and when?
- Tenoke 3mo agohttps://youtube.com/shorts/zQTt8TkcyfU?is=09r7XDqz2w6-Pygu https://youtube.com/shorts/zQTt8TkcyfU?is=09r7XDqz2w6-Pygu You probably wont like the edit but I dont have the timestamp of the original on hand, you can find it.
- typon 3mo agoDinoV3 paper: https://arxiv.org/pdf/2508.10104#page=36 https://arxiv.org/pdf/2508.10104#page=36 "we use a rough estimate of a total 9M GPU hours" From CoreWeave, at current prices (~$2.46/hr spot to ~$6.16/hr on demand) would correspond to $22M–$55M. The dataset is really where the cost is though - they used LVD-1689M - 1.6B images of curated web data from roughly 17B instagram images. This probably cost a huge amount of hours in human annotation, compute for algorithmic filtering, etc and not to mention probably a 20-50 person team working on this model. You might want to change assumptions about how expensive these models are.
- vatsachak 3mo agoThanks for the correction on the order of magnitude for the whole training process. The 9M GPU hours includes the DINO v2 inference used in order to curate the data set. The final training run used like 300000 dollars of compute. Unfortunately we don't know how much RLVR + Agent training costs these companies. I'm just gonna say it's in the hundreds of millions, because they are supposedly making billions of profit on inference yet making billion dollar losses
- ainch 3mo agoYann is a big SSL guy but I don't think he was involved in the original DINO - he's not listed as a co-author or anything.
- vatsachak 3mo agoDINO was created independent of JEPA but uses a similar principle of self supervised learning through minimizing the prediction error of a latent. The difficulty in predicting a latent is so called "collapse"; the embedding neutral network can always output the zero vector and this would predict the output correctly. There are different ways to solve this, DINO uses two different models - a teacher and a student and LeCunn uses an explicit term against collapsing to a single output. Yann mentions DINO in his talks
- chrsw 3mo agoMy main takeaway from LeCun's thesis isn't that you can't build LLMs to do useful things better than the best human, it's that these systems don't learn arbitrary skills efficiently, like humans do. And the question is, why not? 8% on ARC-AGI-3 is amazing for a machine considering how far we've come since digital computers were first built. But it is pretty poor if you're claiming something is well on its way to exhibiting human-like intelligence. Mythos can do some amazing things (I'm assuming, I've never seen it). A young child can learn to control its body without reading any books on dynamical systems and kinematics. Mythos cannot learn to control a humanoid robot after sucking in every piece of data Anthropic can get their hands on.
- Maxatar 3mo agoFalsifying Yann Lecun isn't exactly a priority for anyone seriously working in this space.
- CyberDildonics 3mo agoWhat is this supposed to mean?
- bevekspldnw 3mo ago“Bro” spent most of his career in the wilderness because everybody thought ML/NN/etc were a dead end. I’d not wager against him having at one one more break though architecture before he retires.
- raverbashing 3mo agoNotice how neither him, nor Ilya, nor Mira shipped anything relevant recently It's telling
- CuriouslyC 3mo agoNot sure how Mira gets into the same sentence as Yann and Ilya. As far as the lack of shipping, they're scientists and what we're doing now with LLMs is more "engineering."
- reasonableklout 3mo agoMythos doesn't appear to be on the verified leaderboard for ARC-AGI 3
- Muskwalker 3mo agoOfficially, Fable/Mythos testing was delayed because of Anthropic's data retention policy. Don't know if there's word on them working that out yet. https://x.com/arcprize/status/2064399134099153344 https://x.com/arcprize/status/2064399134099153344
- bansimonw 3mo ago[dead]
- 10xDev 3mo agoSeeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.
- vatsachak 3mo agoI mean, theoretically you can solve every finitary problem with a brute force solution... Richard Sutton specifically states that the search has to be smart. We know that the brain uses recurrent connections and is shallow. I think a lot more money has to go into architecture. Feed Forward transformers can only scale so far
- deleted 3mo ago[deleted]
- altcognito 3mo agoWhile I think this is true, remember as we get more efficient we just decide to scale even bigger. So more GPUs, and more efficient. I agree with the sibling comment, effiency is probably the more important component at this point. We are hitting not just a practical engineering roadblock for scaling with current technology, I think we have definitely hit a financial and logistical roadblock for up scaling with the number of GPUs (on an immediate basis)
- Razengan 3mo ago> We are probably going to need a lot more GPUs. Or a breakthrough in algorithms etc. The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.
- altcognito 3mo ago20 watts for inference AND training!
- aeyes 3mo agoFor intelligence, I expect the next breakthrough to be colocation of memory and compute in the same chip. And we'll need much more of this memory, probably a few petabytes.
- balefulboy 3mo agoit seems the older models were capped at 10kusd for the runs though?
- maxnevermind 3mo agoI'm surprised it is that low. Are not all top AI labs "cheating" and workaround LLMs's low sample efficiency by hiring people to generate more data points - similar problems with answers, so they can train models on those and improve scores? A good benchmark for general intelligence probably should be a complete black box, no sample data given/leaked at all.
- Kuinox 3mo agoOh it's because the bench is lying. You need to pass each level without failing, if you fail a level, it count as you "lost" the minigame. The fact it start to get a score means it managed to get a 100% score on one of the minigame.
- monk_grilla 3mo agoThis is the first I have herd of this benchmark. Can someone explain how it in any way indicates how close we are to "AGI"? Replay of Sol attempting the game: https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9ebcd https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9... It seems a weird and arbitrary challenge for a language model to be expected to perform. It also seems like there are some harness/visual issues even in the first few steps, where it states that it hasn't moved when it clearly has.
- drdrey 3mo agothe problems are general and abstract, domain specific knowledge and memorization don't help. Figuring out the rules, the goal, the controls, and how to solve in a reasonable budget all indicate some level of general ability.
- andai 3mo agoSo basically they're well suited for like, an octopus or a crow? I was thinking about those species earlier in the context of, what does intelligence mean outside of language. The benchmark appears to be testing the same thing. Although I don't know how much transfer there would be between this data set and the kind of situations a crow or an octopus would encounter. Edit: Huh, it's just a Game boy game? I just did a couple of the tasks. It looks like C64 era game to me. Navigating levels. A lot of overlap with animal intelligence then.
- andriy_koval 3mo ago> Can someone explain how it in any way indicates how close we are to "AGI"? I think it is historical name. At some point when benchmarking was very undeveloped, this was targeting abstract reasoning and generalization, hence AGI.
- gertlabs 3mo agoWe have it slightly ahead of Fable in our multi-agent coding evaluations. Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol. Fable scores slightly lower, but with an average solution size of 12.2 KB. Data at https://gertlabs.com/rankings?mode=oneshot_coding https://gertlabs.com/rankings?mode=oneshot_coding
- dgfl 3mo agoThis looks like a good benchmark. Time and time again I keep giving OpenAI models the chance to win me back, but Opus (and Fable especially) just writes more elegant code and is a significantly more productive rubber duck for interactive discussions. I feel vindicated seeing your description of verbose and defensive code, and I’m a bit disappointed that 5.6 Sol’s solution is still >5x longer than the human solution and 2x as verbose as Fable’s. Do you have any insight whether any of that is comments? I wonder why nobody has tried to optimize for actual code size or complexity metric, or at least why I haven’t seen more benchmarks that display this. GPT5.5 just keeps pushing more and more pointless indirection into every function it writes in my main project, it’s borderline negative productivity. P.S. I’d be curious to see Cursor’s composer models in there, they seem to be among the best performing low cost models: https://artificialanalysis.ai/articles/cursor-composer-2-5-coding-agent-index https://artificialanalysis.ai/articles/cursor-composer-2-5-c...
- desterothx 3mo agoReally almost all benchmarks I look at have a cost per task column, which is basically the code size metric if you take an extra step
- dgfl 3mo agoNot at all. The model could (and sometimes should) burn all the money it wants, and then produce a single line of actual production code. Only some things, e.g. full rewrites, have clear cost - LoC scaling. For my usage, I would very much prefer if those $/task were being spent in thinking and experimenting, and the actual output would be as short and maintainable as possible. “maintainability” is a vague target of course, but it’s at least somewhat correlated with code size.
- akoboldfrying 3mo ago> Cost/task: $25.1K Yikes