5 ms·
Does the arc-agi-2 score more than doubling in a .1 release indicate benchmark-maxing? Though i dont know what arc-agi-2 actually tests
by ripbozo 8mo ago
Does the arc-agi-2 score more than doubling in a .1 release indicate benchmark-maxing? Though i dont know what arc-agi-2 actually tests
- blinding-streak 8mo agoI assume all the frontier models are benchmaxxing, so it would make sense
- maxall4 8mo agoTheoretically, you can’t benchmaxx ARC-AGI, but I too am suspect of such a large improvement, especially since the improvement on other benchmarks is not of the same order.
- moffkalast 8mo agohttps://arcprize.org/arc-agi/1/ https://arcprize.org/arc-agi/1/ It's a sort of arbitrary pattern matching thing that can't be trained on in the sense that the MMLU can be, but you can definitely generate billions of examples of this kind of task and train on it, and it will not make the model better on any other task. So in that sense, it absolutely can be. I think it's been harder to solve because it's a visual puzzle, and we know how well today's vision encoders actually work https://arxiv.org/html/2407.06581v1 https://arxiv.org/html/2407.06581v1
- km144 8mo agoThe real question is: Why are people designing benchmarks that, if a model is trained on them, it won't improve the performance of the model at any real-world tasks? Why would anyone care about such benchmarks?
- moffkalast 8mo agoPeople are like typewriter monkeys, if something is possible to make it'll eventually be made.
- boplicity 8mo agoBenchmark maxing could be interpreted as benchmarks actually being a design framework? I'm sure there are pitfalls to this, but it's not necessarily bad either.
- energy123 8mo agoFrancois Chollet accuses the big labs of targeting the benchmark, yes. It is benchmaxxed.
- CamperBob2 8mo agoI don't know what he could mean by that, as the whole idea behind ARC-AGI is to "target the benchmark." Got any links that explain further?
- layer8 8mo agoThe fact that ARC-AGI has public and semi-private in addition to private datasets might explain it: https://arcprize.org/arc-agi/2/#dataset-structure https://arcprize.org/arc-agi/2/#dataset-structure
- tasuki 8mo agoDidn't the same Francois Chollet claim that this was the Real Test of Intelligence? If they target it, perhaps they target... real intelligence?
- ainch 8mo agoHe's always said ARC is a necessary but not sufficient condition for testing intelligence afaik
- energy123 8mo agoHe said in an interview that it doesn't count if it's explicitly targeted, only if a model generalizes to it. He also said that the "real test of intelligence" is being unable to come up with new tests that a human can easily do that the AI can't, not in being able to pass any specific benchmark.
- segmondy 8mo agoHe should have kept it closed.