4 ms·
This is with the caveat that OpenAI uses their own harness for this: > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two set
by aabhay 1mo ago
This is with the caveat that OpenAI uses their own harness for this:
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
- simianwords 1mo agoThis should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
- ActionHank 1mo ago"This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat."
- simianwords 1mo ago> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. > Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches. This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh
- ActionHank 1mo ago"Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results."
- simianwords 1mo ago"I'll twist the words of the author of the benchmark itself to make a point"
- ActionHank 1mo ago"I refuse to see the wall that I am running directly into, because if I see it I will hit it"
- simianwords 1mo agoIf you mean a capability wall, the author of the benchmark says this >We see Astra as a major breakthrough in model intelligence. You think the author of the benchmark is also in the conspiracy
- Readerium 1mo agoIts 62 percent when using a neutral harness. https://arcprize.org/blog/astra https://arcprize.org/blog/astra
- glenstein 1mo agoInteresting both this and Sol got approximately a 37% boost with the custom harness.
- sbinnee 1mo agoYet it is an impressive number. But yeah when you see a number 99 you have doubts. Thanks for the link
- Doohickey-d 1mo agoTheir neutral harness is not very good though, if I read it right, it doesn't preserve the reasoning state between turns. No real harness discards reasoning state like that.
- AmazingTurtle 1mo agoYeah that must be it. OpenAI doesn't want to disclose internal reasoning, that's why thats typically encrypted_content in OpenAI codex session ledgers etc.; leveraging responses API preserves reasoning server side all the way till a final answer is made; so that's very impressive and to me the score that matters.