4 ms·
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-
by scrlk 1mo ago
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/ https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
- kasperni 1mo agoyes it is.
- woah 1mo agoHaven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
- tintor 1mo agoThose people haven't verified their results against the private set: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- andriy_koval 1mo agoAstra also not verified using private set, but on "semi-private" set
- andrewchambers 1mo agoif that is true then why is astra on the official ARC leaderboard now ?
- andriy_koval 1mo agoARC leaderboard has results from semi-private data for frontier models, they have another competition for private data. It is described in their methodology: https://arcprize.org/policy https://arcprize.org/policy It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
- tintor 1mo agoWhere are results for private data? Which LLMs participate on private set? Open weight LLMs only?
- andriy_koval 1mo agoYes, they run competitions once a year amongst open weight models
- enraged_camel 1mo agoYep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
- 10xDev 1mo agoIt is about memory retention. No heavy lifting done on the reasoning side so I hardly see anything misleading here. Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371 https://x.com/fchollet/status/2095598451115614371
- tedsanders 1mo agoOur responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all. ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard https://arcprize.org/leaderboard A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good! (I coauthored the linked blog post)
- GPerson 1mo ago[flagged]