3 ms·
Theoretically, you can’t benchmaxx ARC-AGI, but I too am suspect of such a large improvement, especially since the improvement on other benchmarks is not of the
by maxall4 8mo ago
Theoretically, you can’t benchmaxx ARC-AGI, but I too am suspect of such a large improvement, especially since the improvement on other benchmarks is not of the same order.
- moffkalast 8mo agohttps://arcprize.org/arc-agi/1/ https://arcprize.org/arc-agi/1/ It's a sort of arbitrary pattern matching thing that can't be trained on in the sense that the MMLU can be, but you can definitely generate billions of examples of this kind of task and train on it, and it will not make the model better on any other task. So in that sense, it absolutely can be. I think it's been harder to solve because it's a visual puzzle, and we know how well today's vision encoders actually work https://arxiv.org/html/2407.06581v1 https://arxiv.org/html/2407.06581v1
- km144 8mo agoThe real question is: Why are people designing benchmarks that, if a model is trained on them, it won't improve the performance of the model at any real-world tasks? Why would anyone care about such benchmarks?
- moffkalast 8mo agoPeople are like typewriter monkeys, if something is possible to make it'll eventually be made.