5 ms·
Less than a year to destroy Arc-AGI-2 - wow.
by neilellis 8mo ago
Less than a year to destroy Arc-AGI-2 - wow.
- Davidzheng 8mo agoI unironically believe that arc-agi-3 will have a introduction to solved time of 1 month
- ACCount37 8mo agoNot very likely? ARC-AGI-3 has a nasty combo of spatial reasoning + explore/exploit. It's basically adversarial vs current AIs.
- Davidzheng 8mo agoWe will see at the end of April right? It's more of a guess than a strongly held conviction--but I see models improving rapidly at long horizon tasks so I think it's possible. I think a benchmark which can survive a few months (maybe) would be if it genuinely tested long time-frame continual learning/test-time learning/test-time posttraining (idk honestly the differences b/t these). But i'm not sure how to give such benchmarks. I'm thinking of tasks like learning a language/becoming a master at chess from scratch/becoming a skill artists but where the task is novel enough for the actor to not be anywhere close to proficient at beginning--an example which could be of interest is, here is a robot you control, you can make actions, see results...become proficient at table tennis. Maybe another would be, here is a new video game, obtain the best possible 0% speedrun.
- etyhhgfff 8mo agoThe AGI bar has to be set even higher, yet again.
- red75prime 8mo agoAnd that's the way it should be. We're past the "Look! It can talk! How cute!" stage. AGI should be able to deal with any problem a human can.
- dakolli 8mo agowow solving useless puzzles, such a useful metric!
- esafak 8mo agoHow is spatial reasoning useless??
- saberience 8mo agoIt's a useless meaningless benchmark though, it just got a catchy name, as in, if the models solve this it means they have "AGI", which is clearly rubbish. Arc-AGI score isn't correlated with anything useful.
- jabedude 8mo agohow would we actually objectively measure a model to see if it is AGI if not with benchmarks like arc-AGI?
- WarmWash 8mo agoGive it a prompt like >can u make the progm for helps that with what in need for shpping good cheap products that will display them on screen and have me let the best one to get so that i can quickly hav it at home And get back an automatic coupon code app like the user actually wanted.
- Legend2440 8mo agoIt's correlated with the ability to solve logic puzzles. It's also interesting because it's very very hard for base LLMs, even if you try to "cheat" by training on millions of ARC-like problems. Reasoning LLMs show genuine improvement on this type of problem.
- HDThoreaun 8mo agoARC-AGI 2 is an IQ test. IQ tests have been shown over and over to have predictive power in humans. People who score well on them tend to be more successful
- fsh 8mo agoIQ tests only work if the participants haven't trained for them. If they do similar tests a few times in a row, scores increase a lot. Current LLMs are hyper-optimized for the particular types of puzzles contained in popular "benchmarks".
- XCSme 8mo agoBut why only a +0.5% increase for MMMU-Pro?
- kenjackson 8mo agoEveryone is already at 80% for that one. Crazy that we were just at 50% with GPT-4o not that long ago.
- XCSme 8mo agoBut 80% sounds far from good enough, that's 20% error rate, unusable in autonomous tasks. Why stop at 80%? If we aim for AGI, it should 100% any benchmark we give.
- kenjackson 8mo agoAre humans 100%?
- XCSme 8mo agoIf they are knowledgeable enough and pay attention, yes. Also, if they are given enough time for the task. But the idea of automation is to make a lot fewer mistakes than a human, not just to do things faster and worse.
- kenjackson 8mo agoActually faster and worse is a very common characterization of a LOT of automation.
- XCSme 8mo agoThat's true. The problem is that if the automation breaks at any point, the entire system fails. And programming automations are extremely sensitive to minor errors (i.e. a missing semicolon). AI does have an interesting feature though, it tends to self-healing in a way, when given tools access and a feedback loop. The only problem is that self-healing can incorrectly heal errors, then the final reault will be wrong in hard-to-detect ways. So the more wuch hidden bugs there are, the nore unexpectedly the automations will perform. I still don't trust current AI for any tasks more than data parsing/classification/translation and very strict tool usage. I don't beleive in the full-assistant/clawdbot usage safety and reliability at this time (it might be good enough but the end of the year, but then the SWE bench should be at 100%).
- modeless 8mo agoIt's still useful as a benchmark of cost/efficiency.