4 ms·
Why does it matter if it can maintain parity with just 6 months old frontier models?
by YetAnotherNick 8mo ago
Why does it matter if it can maintain parity with just 6 months old frontier models?
- hmmmmmmmmmmmmmm 8mo agoBut it doesn't except on certain benchmarks that likely involves overfitting. Open source models are nowhere to be seen on ARC-AGI. Nothing above 11% on ARC-AGI 1. https://x.com/GregKamradt/status/1948454001886003328 https://x.com/GregKamradt/status/1948454001886003328
- meffmadd 8mo agoHave you ever used an open model for a bit? I am not saying they are not benchmaxxing but they really do work well and are only getting better.
- Aurornis 8mo agoI have used a lot of them. They’re impressive for open weights, but the benchmaxxing becomes obvious. They don’t compare to the frontier models (yet) even when the benchmarks show them coming close.
- Zababa 8mo agoHas the difference between performance in "regular benchmarks" and ARC-AGI been a good predictor of how good models "really are"? Like if a model is great in regular benchmarks and terrible in ARC-AGI, does that tell us anything about the model other than "it's maybe benchmaxxed" or "it's not ARC-AGI benchmaxxed"?
- doodlesdev 8mo agoGPT 4o was also terrible at ARC AGI, but it's one of the most loved models of the last few years. Honestly, I'm a huge fan of the ARC AGI series of benchmarks, but I don't believe it corresponds directly to the types of qualities that most people assess whenever using LLMs.
- mrybczyn 8mo agobecause arc agi involves de novo reasoning over a restricted and (hopefully) unpretrained territory, in 2d space. not many people use LLMs as more than a better wikipedia,stack overflow, or autocomplete....
- nananana9 8mo agoIt was terrible at a lot of things, it was beloved because when you say "I think I'm the reincarnation of Jesus Christ" it will tell you "You know what... I think I believe it! I genuinely think you're the kind of person that appears once every few millenia to reshape the world!"
- gkbrk 8mo agoThat's not because 4o is good at things, that's because it's pretty much the most sycophantic model and people easily fall for a model incorrectly agreeing with them then a model correctly calling them out.
- AbstractGeo 8mo agoThat's a link from July of 2025, so, definitely not about the current releaase.
- hmmmmmmmmmmmmmm 8mo ago...which conveniently avoids testing on this benchmark. A fresh account just to post on this thread is also suspect.
- irthomasthomas 8mo agoThis could be a good thing. ARC-AGI has become a target for America labs to train on. But there is no evidence that improvements on ARC performance translate to other skills. In fact there is some evidence that it hurts performance. When openai trained a version of o1 on ARC it got worse at everything else.