2 ms·
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010 https://x.com/PrimeIntellect/status/2085087000764
by stared 2mo ago
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010 https://x.com/PrimeIntellect/status/2085087000764568010.
I am curious - how does it fare for other benchmarks, or everyday programming?
- tintor 2mo agoPrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- stared 2mo agoGood to know! Is it that it wasn't accepted yet, or are there issues with how it was run?
- noahbp 2mo agoIt’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers. There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.
- andriy_koval 2mo agoleaderboard likely has results from "semi-private" dataset, and graph above likely from public dataset, so it can be easily overfit.