10 ms·
Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-deep-think/ https://blog.google
by lukebechtel 8mo ago
Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6)
Wow.
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-deep-think/ https://blog.google/innovation-and-ai/models-and-research/ge...
- karmasimida 8mo agoIt is over
- baal80spam 8mo agoI for one welcome our new AI overlords.
- mnicky 8mo agoWell, fair comparison would be with GPT-5.x Pro, which is the same class of a model as Gemini Deep Think.
- nubg 8mo agoWeren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.
- verdverm 8mo agoHere's a good thread over 1+ month, as each model comes out https://bsky.app/profile/pekka.bsky.social/post/3meokmizvt22v https://bsky.app/profile/pekka.bsky.social/post/3meokmizvt22... tl;dr - Pekka says Arc-AGI-2 is now toast as a benchmark
- Aperocky 8mo agoIf you look at the problem space it is easy to see why it's toast, maybe there's intelligence in there, but hardly general.
- verdverm 8mo agothe best way I've seen this describes is "spikey" intelligence, really good at some points, those make the spikes humans are the same way, we all have a unique spike pattern, interests and talents ai are effectively the same spikes across instances, if simplified. I could argue self driving vs chatbots vs world models vs game playing might constitute enough variation. I would not say the same of Gemini vs Claude vs ... (instances), that's where I see "spikey clones"
- Aperocky 8mo agoYou can get more spiky with AIs, whereas with human brain we are more hard wired. So maybe we are forced to be more balanced and general whereas AI don't have to.
- verdverm 8mo agoI suspect the non-spikey part is the more interesting comparison Why is it so easy for me to open the car door, get in, close the door, buckle up. You can do this in the dark and without looking. There are an infinite number of little things like this you think zero about, take near zero energy, yet which are extremely hard for Ai
- gowld 8mo agoYou are asking a robotics question, not an AI question. Robotics is more and less than AI. Boston Dynamics robots are getting quite near your benchmark.
- idiotsecant 8mo agoBoston dynamics is missing just about all the degrees of freedom involved in the scenario op mentions.
- pixl97 8mo ago>Why is it so easy for me to open the car door Because this part of your brain has been optimized for hundreds of millions of years. It's been around a long ass time and takes an amazingly low amount of energy to do these things. On the other hand the 'thinking' part of your brain, that is your higher intelligence is very new to evolution. It's expensive to run. It's problematic when giving birth. It's really slow with things like numbers, heck a tiny calculator and whip your butt in adding. There's a term for this, but I can't think of it at the moment.
- fishpham 8mo agoYes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)
- jstummbillig 8mo agoCould it also be that the models are just a lot better than a year ago?
- bigbadfeline 8mo ago> Could it also be that the models are just a lot better than a year ago? No, the proof is in the pudding. After AI we're having higher prices, higher deficits and lower standard of living. Electricity, computers and everything else costs more. "Doing better" can only be justified by that real benchmark. If Gemini 3 DT was better we would have falling prices of electricity and everything else at least until they get to pre-2019 levels.
- ctoth 8mo ago> If Gemini 3 DT was better we would have falling prices of electricity and everything else at least Man, I've seen some maintenance folks down on the field before working on them goalposts but I'm pretty sure this is the first time I saw aliens from another Universe literally teleport in, grab the goalposts, and teleport out.
- WarmWash 8mo agoYou might call me crazy, but at least in 2024, consumers spent ~1% less of their income on expenses than 2019[2], which suggests that 2024 is more affordable than 2019. This is from the BLS consumer survey report released in dec[1] [1]https://www.bls.gov/news.release/cesan.nr0.htm https://www.bls.gov/news.release/cesan.nr0.htm [2]https://www.bls.gov/opub/reports/consumer-expenditures/2019/ https://www.bls.gov/opub/reports/consumer-expenditures/2019/ Prices are never going back to 2019 numbers though
- modeless 8mo agoFrançois Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI.
- hmmmmmmmmmmmmmm 8mo agoI don't think the creator believes ARC3 can't be solved but rather that it can't be solved "efficiently" and >$13 per task for ARC2 is certainly not efficient. But at this rate, the people who talk about the goal posts shifting even once we achieve AGI may end up correct, though I don't think this benchmark is particularly great either.
- beklein 8mo agohttps://x.com/fchollet/status/2022036543582638517 https://x.com/fchollet/status/2022036543582638517
- joelthelion 8mo agoDo opus 4.6 or gemini deep think really use test time adaptation ? How does it work in practice?
- mapontosevenths 8mo ago> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests they used to give black people to prevent them from voting. We dont ask nearly so much proof from a human, we take their word for it. On the few occasions we did ask for proof it inevitably led to horrific abuse. Edit: The average human tested scores 60%. So the machines are already smarter on an individual basis than the average human.
- aeyes 8mo agohttps://arcprize.org/leaderboard https://arcprize.org/leaderboard $13.62 per task - so we need another 5-10 years for the price to run this to become reasonable? But the real question is if they just fit the model to the benchmark.
- igravious 8mo agoThat's not a long time in the grand scheme of things.
- throwup238 8mo agoSpeak for yourself. Five years is a long time to wait for my plans of world domination.
- amelius 8mo agoYes, you better hurry.
- tasuki 8mo agoThis concerns me actually. With enough people (n>=2) wanting to achieve world domination, we have a problem.
- gowld 8mo agon = 2 is Pinky and the Brain.
- antonvs 8mo agoI'm convinced that a substantial fraction of current tech CEOs were unwittingly programmed as children by that show.
- throwup238 8mo agoIt’s not that I want to achieve world domination (imagine how much work that would be!), it’s just that it’s the inevitable path for AI and I’d rather it be me than then next shmuck with a Claude Max subscription.
- saberience 8mo agoArc-AGI (and Arc-AGI-2) is the most overhyped benchmark around though. It's completely misnamed. It should be called useless visual puzzle benchmark 2. It's a visual puzzle, making it way easier for humans than for models trained on text firstly. Secondly, it's not really that obvious or easy for humans to solve themselves! So the idea that if an AI can solve "Arc-AGI" or "Arc-AGI-2" it's super smart or even "AGI" is frankly ridiculous. It's a puzzle that means nothing basically, other than the models can now solve "Arc-AGI"
- CuriouslyC 8mo agoThe puzzles are calibrated for human solve rates, but otherwise I agree.
- saberience 8mo agoMy two elderly parents cannot solve Arc-AGI puzzles, but can manage to navigate the physical world, their house, garden, make meals, clean the house, use the TV, etc. I would say they do have "general intelligence", so whatever Arc-AGI is "solving" it's definitely not "AGI"
- hmmmmmmmmmmmmmm 8mo agoYou are confusing fluid intelligence with crystallised intelligence.
- casey2 8mo agoI think you are making that confusion. Any robotic system in the place of his parents would fail with a few hours. There are more novel tasks in a day than ARC provides.
- hmmmmmmmmmmmmmm 8mo agoChildren have great levels of fluid intelligence, that's how they are able to learn to quickly navigate in a world that they are still very new to. Seniors with decreasing capacity increasingly rely on crystallised intelligence, that's why they can still perform tasks like driving a car but can fail at completely novel tasks, sometimes even using a smartphone if they have not used one before.
- mNovak 8mo agoI'm excited for the big jump in ARC-AGI scores from recent models, but no one should think for a second this is some leap in "general intelligence". I joke to myself that the G in ARC-AGI is "graphical". I think what's held back models on ARC-AGI is their terrible spatial reasoning, and I'm guessing that's what the recent models have cracked. Looking forward to ARC-AGI 3, which focuses on trial and error and exploring a set of constraints via games.
- throw310822 8mo agoThe average ARC AGI 2 score for a single human is around 60%. "100% of tasks have been solved by at least 2 humans (many by more) in under 2 attempts. The average test-taker score was 60%." https://arcprize.org/arc-agi/2/ https://arcprize.org/arc-agi/2/
- modeless 8mo agoWorth keeping in mind that in this case the test takers were random members of the general public. The score of e.g. people with bachelor's degrees in science and engineering would be significantly higher.
- throw310822 8mo agoRandom members of the public = average human beings. I thought those were already classified as General Intelligences.
- thesmtsolver2 8mo agoAverage human beings with average human problems.
- imiric 8mo agoWhat is the point of comparing performance of these tools to humans? Machines have been able to accomplish specific tasks better than humans since the industrial revolution. Yet we don't ascribe intelligence to a calculator. None of these benchmarks prove these tools are intelligent, let alone generally intelligent. The hubris and grift are exhausting.
- raincole 8mo agoEven before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro playthroughs in text form online I think it's a much stronger signal of its 'generalness' than ARC-AGI. By the way, Deepseek can't play Balatro at all. [0]: https://balatrobench.com/ https://balatrobench.com/
- winstonp 8mo agoDeepSeek hasn't been SotA in at least 12 calendar months, which might as well be a decade in LLM years
- cachius 8mo agoWhat about Kimi and GLM?
- zozbot234 8mo agoThese are well behind the general state of the art (1yr or so), though they're arguably the best openly-available models.
- tgrowazay 8mo agoAccording to artificial analysis ranking, GLM-5 is at #4 after Claude Opus 4.5, GPT-5.2-xhigh and Claude Opus 4.6 .
- epolanski 8mo agoIdk man, GLM 5 in my tests matches opus 4.5 which is what, two months old?
- wahnfrieden 8mo ago
- deleted 8mo ago[deleted]
- culi 8mo agoYes but with a significant (logarithmic) increase in cost per task. The ARC-AGI site is less misleading and shows how GPT and Claude are not actually far behind https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- robertwt7 8mo agoI’m surprised that gemini 3 pro is so low at 31.1% though compared to opus 4.6 and gpt 5.2. This is a great achievement but its only available to ultra subscribers unfortunately
- chillfox 8mo agoAt $13.62 per task it's practically unusable for agent tasks due to the cost. I found that anything over $2/task on Arc-AGI-2 ends up being way to much for use in coding agents.
- thefounder 8mo agoAm I the only one that can’t find Gemini useful except if you want something cheap? I don’t get what was the whole code red about or all that PR. To me I see no reason to use Gemini instead of of GPT and Anthropic combo. I should add that I’ve tried it as chat bot, coding through copilot and also as part of a multi model prompt generation. Gemini was always the worst by a big margin. I see some people saying it is smarter but it doesn’t seem smart at all.
- pell 8mo agoI find the quality is not consistent at all and of all the LLMs I use Gemini is the one most likely to just verge off and ignore my instructions.
- Foobar8568 8mo agoSame, as far as I am concerned, Gemini is optimized for benchmarks. I mean last week it insisted suddenly on two consecutive prompts that my code was in python. It was in rust.
- Nathanba 8mo agoYou are not the only one, it's to the point where I think that these benchmark results must be faked somehow because it doesn't match my reality at all.
- viking123 8mo agoIt's garbage really, cannot get how they get so high in benchmarks.
- mileshilles 8mo agomaybe it depends on the usage, but in my experience most of the times the Gemini produces much better results for coding, especially for optimization parts. The results that were produced by Claude wasn't even near that of Gemini. But again, depends on the task I think.
- nprateem 8mo agoYeah it's pretty shit compared to Opus
- fzeindl 8mo agoI read somewhere that Google will ultimately always produce the best LLMs, since "good AI" relies on massive amounts of data and Google owns the most data. Is that a based assumption?
- astrange 8mo agoNo.
- SV_BubbleTime 8mo agoCorrect. Great output is a good model with good context… at the right time. Google isn’t guaranteed any of these.
- whiplash451 8mo agoWe can really look at it both ways. It is actually concerning that a model that won IMO last summer would still fail 15% of ARC AGI 2.
- emp17344 8mo agoI mean, remember when ARC 1 was basically solved, and then ARC 2 (which is even easier for humans) came out, and all of the sudden the same models that were doing well on ARC 1 couldn’t even get 5% on ARC 2? Not convinced this isn’t data leakage.