4 ms·
ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard https://arcprize.org/leaderboard
by tintor 1mo ago
ARC AGI-3 saturated by Astra!
https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- IshKebab 1mo agoLook at those costs!
- schaefer 1mo agoRight? Between $18k-40k to run a benchmark.
- andriy_koval 1mo agoI think it could indicate that "semi-private" dataset likely leaked to their training data.
- IshKebab 1mo agoIt says "Provider Adapter" so presumably they put some manual work in to make this work.
- xpct 1mo agoA dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems. Still, probably not that much compared to employees targeting it.
- minimaxir 1mo agoARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra https://arcprize.org/blog/astra tl;dr it's 62% when apples-to-apples to other models, which is still notable.
- ciefa 1mo agoWoah, that is a crazy interesting read!
- debazel 1mo agoARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
- Readerium 1mo agosaturated before (higher degree) AGI-2
- vb-8448 1mo agoBut scored less on V2 and V1 ... too much overfitting?
- zem 1mo agohttps://mvakde.github.io/blog/44-on-arc-1/ https://mvakde.github.io/blog/44-on-arc-1/ makes a good case that all the performance on the arc agi tests is overfitting, based on the fact that v1 performance did not translate directly to v2 performance
- XCSme 1mo agoThe no-reasoning version scores 35% while the low reasoning one scores 17%? What?
- silver_sun 1mo agoIt's simulating the Dunning-Kruger effect.
- dudeinhawaii 1mo agoI suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".
- XCSme 1mo agoIt says (none), which usually means reasoning disabled. I would be surprised if none = use default reasoning
- tom2026hn 1mo agoAGI means that the answers to these benchmarks are now accessible, and the defense capabilities of the testing organizations are now negligible.