4 ms·
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
by woah 1mo ago
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
- tintor 1mo agoThose people haven't verified their results against the private set: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- andriy_koval 1mo agoAstra also not verified using private set, but on "semi-private" set
- andrewchambers 1mo agoif that is true then why is astra on the official ARC leaderboard now ?
- andriy_koval 1mo agoARC leaderboard has results from semi-private data for frontier models, they have another competition for private data. It is described in their methodology: https://arcprize.org/policy https://arcprize.org/policy It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
- tintor 1mo agoWhere are results for private data? Which LLMs participate on private set? Open weight LLMs only?
- andriy_koval 1mo agoYes, they run competitions once a year amongst open weight models