4 ms·
Good to know! Is it that it wasn't accepted yet, or are there issues with how it was run?
by stared 2mo ago
Good to know!
Is it that it wasn't accepted yet, or are there issues with how it was run?
- noahbp 2mo agoIt’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers. There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.
- andriy_koval 2mo agoleaderboard likely has results from "semi-private" dataset, and graph above likely from public dataset, so it can be easily overfit.