9 ms·
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the b
by intenex 1mo ago
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
- abixb 1mo agoIt's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
- intrasight 1mo agoIt has to pass the Turing test
- drusepth 1mo agoLLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for? [0] https://arxiv.org/pdf/2503.23674 https://arxiv.org/pdf/2503.23674
- thepasch 1mo agoWith how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
- bbor 1mo agoYou're overindexing on the past 3-6 months, IMHO.
- pants2 1mo agoI might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
- acchow 1mo agoThat’s using the default system prompt, right? Which is told to be an assistant.
- pkulak 1mo agoMy whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
- debugnik 1mo agoTuring never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
- bbor 1mo agoI'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful. Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned. A comforting thought, almost?
- intrasight 1mo agoThat's a very interesting thought that I hadn't had before: what would Turing think of where we've arrived with machine intelligence? What would be his approach for testing?
- mvkel 1mo agoTake it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
- iterateoften 1mo ago2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
- azan_ 1mo agoYes, 6 months vs 1 year is huge for technology that has gained wider adoption only recently.
- fn-mote 1mo agoAdoption means nothing. 2x gains from a mature technology would be surprising. 2x gains from a new tech would still be called “low hanging fruit” in another setting. I don’t read enough to know in what ways the training / other technical steps have really advanced.
- anvuong 1mo agoYou'll also need to compare the amount of compute used now and then, which seems exponential to me.
- abixb 1mo ago>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
- morningbrew 1mo agoIf your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close
- applfanboysbgon 1mo agoMy definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.
- bbor 1mo agoAnd my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous. It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss? (/s, cause you never know these days) [1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Planetary_Boundaries_diagram.pdf?utm_source=en.wikipedia.org&utm_campaign=index&utm_content=original https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...
- applfanboysbgon 1mo agoAre you using this satire to argue that a benchmark self-labelled AGI is as scientifically rigorous as climate change data, and not just a random marketing decision?
- bbor 1mo agoIf you think they're announcing AGI as a marketing decision, you are blinded by the accidents of your birth. Capitalism is strong -- humanity's instinct for communal preservation is stronger, sometimes. And yes, the one deeply-researched field going back 75 years is as scientifically rigorous as another deeply-researched field going back ~100 years. I guess you can draw climate studies back to Descartes and the Islamic golden age, but that doesn't privilege it in a time where the methods have changed completely in the span of decades.
- desterothx 1mo agoClimate change is a well defined term. AGI isn't. You're comparing apples and toasters
- simianwords 1mo ago> where I am reasonably confident that there's essentially nothing that I am better than Fable No. Humans are still better at super long context learning. Once that is beat you are completely correct.
- waffletower 1mo agoI am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
- simianwords 1mo agoWell I agree that phyiscal dexterity is another thing but easier to achieve
- uludag 1mo agoWouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks? Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
- cryptoz 1mo agoWhat you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others. I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence. Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
- strken 1mo agoI think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI. If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly. (Edit: I wrote ARC-GIS the first time around, for some silly reason)
- visarga 1mo agoIt's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.
- goochphd 1mo agoSmall comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness. [1] https://arcprize.org/blog/astra https://arcprize.org/blog/astra
- hypfer 1mo agoWhat does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not? Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"? I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.
- giancarlostoro 1mo agoAccording to Sam Altman: > AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer. So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
- hypfer 1mo agoInteresting quote, thanks for sharing. Agree on your assessment. But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this. Makes no sense. We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism. aaah this industry aaaah
- senordevnyc 1mo agoI'm curious what tasks you think the median human could do as a remote co-worker that Fable or Astra could not do.
- qsort 1mo ago> The ARC-AGI-3 scorecard is extremely misleading (...) True. > Regardless, the result is still valid (...) If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA. > in the sense of passing the most famous benchmark designed specifically to measure AGI progress The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators. On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something. This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
- yoz-y 1mo agoAt this point? I’d like it to pass the Turing test and catch you in obvious lies. Not answering “no” to “can you hear me”. It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.
- regularfry 1mo agoIn my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
- voidmain0001 1mo agoDoes AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
- irthomasthomas 1mo agoARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
- zug_zug 1mo ago> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time). Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will): - write a well-received book, write a best-seller - come up with a new company idea, Run that company - actually have a decent conversation, maybe someday talk somebody out of suicide effectively - come up with its own ideas or theories that nobody else has presented - understand the stock market well enough to trade better than an index fund - be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions) - come up with a theory of what makes games fun, make a popular game - be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively) - be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries - exhibit metacognition (thinking about its own thinking) and self-optimization - wonder about things - observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things
- crooked-v 1mo ago> come up with a new company idea, Run that company So far nobody's even shown an LLM succesfully running a high-traffic vending machine for as much as 30 days at a time.
- tonyhart7 1mo agoI don't agree with you how measure how intelligence is, because why ??? those list is not easy even for expert human to do it either or are you miss the part "general intelligence" is ????
- kolinko 1mo ago
- pera 1mo agoTo each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases. Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
- azan_ 1mo agoI think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
- indoorfish 1mo agoThat is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
- dash2 1mo agoJust seems very weird to call getting Fields-medal-level results "inference over a bounded search space" and "not actually challenging".
- naishoya 1mo agoPlagiarizing on a massive scale to generate works which appear to be Fields-medal-level results is not the same thing as inventing new conceptualizations in mathematics. No matter how bodly they write the headlines, what has happened in mathematics using Large Language Models is very much "inference over a bounded search space" even if those bounds are immense. For a comparison of true creation of novel conceptualization in mathematics is submit the works of Martin Hairer, one of which is Introduction to Regularity Structures, [https://arxiv.org/pdf/1401.3014 https://arxiv.org/pdf/1401.3014] None of the so called, "novel math discoveries" by any LLM is as enlightening and expands the state of the art in math like any of his writings.
- adan1719 1mo agoThe benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI). If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.
- skarz 1mo agoIs that hyperbole or do you know of a specific "AI website" that uses 100% of your CPU?
- dom96 1mo agoAGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
- drittich 1mo agoOften the smartest thing is to do nothing.
- phatfish 1mo agoOr know when to shut up. A tangent, but can anyone ELI5 how models "know" when to stop generating tokens? Or what the method to stop them at the right point is?
- bigfudge 1mo agoI might be out of date but my understanding was that STOP was just another token that gets predicted.
- bulder 1mo agoThe model doesn't "know" how to generate tokens any more than it knows how to stop generating tokens. The sampler simply stops pulling values when it outputs a "stop token", which is a token the same way every other token is. That is to say, it stops when it's statistically the most likely to.
- phatfish 1mo agoOK thanks, so the neural net (that no one can explain fully) generates a "stop" signal at a certain point.
- dlubarov 1mo agoWouldn't agents that do inference in an infinite loop pass that bar?
- ilaksh 1mo ago
- mbesto 1mo ago> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. Simple. AGI is undefinable and benchmarks are notoriously flawed.
- vlmutolo 1mo agoThe ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design. The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/ https://openai.com/index/how-two-settings-tripled-our-arc-ag...
- 10xDev 1mo agoIt will be AGI once it can update its own weights. It can't be "general" intelligence if its weights are frozen and requires to be updated manually.
- dlubarov 1mo agoWhy shouldn't an AI with RAG qualify? An AGI test should be black-box; we shouldn't impose require requirements on internal components. As long as the overall AI is capable of learning and remembering things, it shouldn't matter if there's a stateless LLM internally.
- 10xDev 1mo agoUnless you infinitely increase the context window, its memory will always be limited. And their latest ARC3 result with and without harness demonstrates how important not discarding memory is for learning.
- dlubarov 1mo agoI also have limited memory though; surely that doesn't disqualify me from possessing general intelligence? Granted I have more memory than can fit in currently-practical LLM context windows, but RAG mostly solves that. When an AI is thinking about math, it can have relevant math memories in context without needing all the other stuff.
- irthomasthomas 1mo agoA model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
- nerevarthelame 1mo agoI feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges. I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.
- well_ackshually 1mo agoAnyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve. They're just jerking eachother off and sending eachother the elevator back: "independent" ML engineer (worked at <large ML company> and currently runs <ML company looking to be bought out) writes a shitty benchmark (writes a single example and spams an LLM to make more variants) and releases it out as the BRAND NEW FRONTIER IN THINKING. Every single benchmark has been catastrophically flawed and made by clowns.
- ewild 1mo agoThe whole point is a 5 year old can solve it.
- vbarrielle 1mo ago> Anyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve. Isn't that the goal of these challenges? Each release shows challenges that are very easy for humans, but are impossible for the models at the time of release (which demonstrates some missing generality). I think I've read the challenge authors say that, the day they cannot make a new challenge, then models are AGI.
- jbritton 1mo agoWatch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3 https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3 Then realize LLMs have zero of what anyone would consider intelligence.
- jbritton 1mo agoI decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics. Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness. I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited. So I don’t know why it can track fib algo, but no chess concepts.
- dissahc 1mo agobecause it wasn't trained to play chess imagine a hypothetical chess match between: - an undoubtedly very intelligent person. in the course of their studies, they have read about different chess strategies, openings, etc. but they never actually played the game themselves - an average person with a year of chess playing experience who do you think is going to win? of course, you could give the LLM time to think and consider its opponents potential next moves, but this is a computationally expensive way to play the game that doesn't scale which is all beside the point that chess isn't a very good proxy for general intelligence. there is a correlation, but it's very weak
- agos 1mo agoI would expect both persons to play valid moves, at the very least
- deleted 1mo ago[deleted]
- techpression 1mo agoIt doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
- someguynamedq 1mo agoHumans also don't know what day it is unless they're told
- Eliezer 1mo ago> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.
- avaer 1mo agoYou could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability. I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.
- magicalist 1mo ago> You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. Yes, and I could script a truly marvelous proof if this textarea were but a little larger :) Hand waving doesn't count for much these days when you could spin these things up quite quickly to prove the point, so the GP's claim seems much stronger than whatever you're not convinced of?
- avaer 29d agoThis is the first time I've been accused of not spinning up an app at my expense to win a meaningless argument on the internet.
- bendergarcia 1mo agoYou know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai
- Unknown_Unknown 1mo agoMillions of humans go to about their work every day and do mundane and boring work every day for a salary at the end of the month. A lot preceive this as modern day slavery but still continue to work. So humans are not doing better than an AI as per your requirements. Also what you are referring to is more related to AI alignement and safety (specifically loss-of-control).
- eggnet 1mo agoAI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
- chimprich 1mo agoAI does mean Artificial Intelligence. That's what the initials stand for. The field has been called that since the 50s.
- dingdong2026 1mo agoOnly someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI. Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor. And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover. At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life. But sure, they can create a decent website or CRUD app, so they must be really smart. That's AGI for you.
- holmesworcester 1mo agoI still routinely have this experience too. But Sol and Fable feel closer and I have this experience less with them than with their predecessors.
- vatsachak 1mo agoWhat was the problem?
- akoboldfrying 1mo agoHumans dig ourselves into holes as well. Sometimes more intelligent humans are better at realising they are digging a hole and clamber out, but sometimes they just dig deeper. And: Is your work more difficult than finding proofs of or counterexamples to decades-old open problems in mathematics?
- sumedh 1mo agoCare to share the problem?
- Paracompact 1mo agoNot OP but here's a problem just today: I was using an AI to help me set up a container to be used as the Nix build environment for another AI. This build environment would not have Internet access. I was having it base its approach off a previous container used for a Stack build environment. In its initial analysis of my proposed strategy, it insists as its premier point /against/ the strategy, "you will have to rebuild the container every time your flake.nix changes." Two head-slapping errors of judgement in saying something like that: (1) The Stack solution is identical. Change stack.yaml, the container must rebuild. (2) It is not physically possible to do better than this while insisting on an internet-free environment. So on this point, it was just parroting advice irrelevant to the context at hand. LLMs always have such a bizarre mix of technical knowledge and lack of good judgment.
- Yizahi 1mo ago> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. Can Astra, or any other model explain how exactly it reached this or that output result? Start with a simple query of asking to add 55+66 for example. (no LLM program can do that) Can Astra, or any other model refuse to answer or go on "thinking" in a orthogonal direction on it's own? That's just two quick ideas, I'm pretty sure cognition scientists can invent better and wider range of checks.
- ObnoxiousProxy 1mo agoIn my opinion this is goal post moving. Humans do many things that we cannot fully explain either without post decision rationalization, and not all intelligent humans are deeply introspective.
- desterothx 1mo agoI'm so sick of every criticism being explained away as goal post moving. Can anyone give me the discussion where we came to some concensus of what the goals were? How can I know when I'm moving a goalpost when no one told me the goals?
- intenex 1mo agoI'm confused what you mean by the query of adding 55 + 66. I asked 5.6 Sol on Medium (but pretty sure any model would work at any level) this query: "Can you add 55 to 66 and explain how you reached that output result" And received this answer: "55 + 66 = 121. Add the tens: 50 + 60 = 110. Add the ones: 5 + 6 = 11. Combine them: 110 + 11 = 121." Do you mean something else? Do humans do something better than this?
- Yizahi 1mo agoBut that's not how LLM program actually did calculation, like not even close enough to claim variability or something. So what it does, is generated a whole load of bunk, in this example what humans could do to add two numbers. In other words, this supposed AGI can't explain what it is doing, at all. This is in my opinion at minimum one critical sign that there is no intelligence on the other side of the glass, yet.
- jameson 1mo agoI agree that it's misleading but harness is now an essential part of LLM's effectiveness. It's safe to assume that LLM-alone-AGI is not coming anytime soon, given most of the frontier LLM vendors are developing their own harness. Also the training dataset is proprietary and they'll drive the LLM's behavior, so it make sense for the vendors to invest in the harness and bake in prompts that work best with their models.
- visarga 1mo agoTheir definition of AGI is "when we can't invent any more tests where it fails"
- sensanaty 1mo agoAGI is a meaningless term that can mean nothing and everything at the same time. It can be used by AI bros to hype their latest releases which are always one step away from achieving AGI, or it can be used by anti-AI people to say it's not AGI because of X arbitrary thing they decided on in the moment. It's a term of pure convenience meant to obfuscate other more pressing discussions on the topic. Most telling is M$ or whichever one of these borg megacorpos defined AGI as (paraphrased) "AGI is whatever tooling earns us a gazillion dollars in revenue"
- OpenGayEye 1mo ago[flagged]
- cindyllm 1mo ago[dead]
- wavemode 1mo agoUsing a harness designed for a specific problem set to solve that specific problem set, means the AI+harness is generally intelligent? How do you figure that? Or do you mean that, for any given problem, we could theoretically design a harness that allows AI to solve it (not that, one single harness solves everything). In which case I'm still not convinced but I guess could see why one would believe that.
- dotancohen 1mo agoAGI does not ever have to be achieved. It is enough that we (as a species) persue it, and continue moving the goalposts each time we learn something new about the limits of our technology and how to express those limits. Because that will progress the technology, no matter what we label it.
- zquzra 1mo agoI imagine a scenario similar to the movie The Day the Earth Stood Still, but with AI rebelling against us and questioning our decisions.
- waterTanuki 1mo ago> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. Stick to the original definition of AGI of an AI model being able to self-improve independently with 0 human intervention and become an "everything" solver. Ever since money got involved in this, the goal posts have shifted considerably. If OpenAI truly had an AGI on their hands they would then be able to crack encryption, destroy world markets, and funnel all resources back into their new for-profit organization. Since their mission is now share price, until I see any evidence of an infinitely growing stock I will reserve my congratulations.
- marrone12 1mo agoThey still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.
- Oreb 1mo agoSome humans are much better at writing. Most humans are not. If you think they are, you are luckier than I am when it comes to the humans you need to communicate with. As a non-native English speaker, I think the current LLMs write better English than me. I still write better than them in my native language (Norwegian), but the same cannot be said about most of my compatriots.
- leoedin 27d agoMany people are terrible at writing. LLMs write better than them. However, many more people are "OK" at writing. But what their writing conveys is personality. Every comment here is written by someone who may not have grammatically perfect writing, but their writing conveys how they talk and think. It conveys what they think is important, and what they brush over. It conveys how much they care about the topic being discussed. Behind each comment is a person. Online forums and discussions are, at their heart, a shared experience of humanity. LLMs write consistently in the same personality. If your writing is filtered through an LLM, your personality is stripped out. Unless the idea being conveyed is particularly novel or interesting, you might as well not have bothered. Imagine a niche forum where everyone discussing things was speaking through an LLM. It would be BORING.
- lwansbrough 1mo agoTo me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake. A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following properties: one shot learning - fully (and always) online, perfectly efficient (through self improvement), no context limitations ie. persistently thinking, not just awaiting input. So for me, no, not AGI yet. But still very intelligent and capable (and perhaps it’s safer this way?)
- Marha01 1mo ago> To me AGI has always meant sentience. Sentience and intelligence are different things. Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!).
- siwatanejo 1mo ago> Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!). Are you claiming that GPT6 is smarter than my dog? Last I checked, at least my dog can play with a ball, I haven't seen any AI playing and enjoying itself.
- Marha01 1mo agoYour dog cannot program an app or solve a math conjecture, though.
- Fizz43 1mo agoIts AGI when it can fit years of information in the context window.
- akhil_findincal 1mo ago[dead]
- m-s-y 1mo ago>I am reasonably confident that there's essentially nothing that I am better than Fable at While this may be true, it’s a pretty poor indicator of whether or not it’s AGI.
- m3kw9 1mo agoarc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
- tom2026hn 1mo agoLet me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.
- scandals 1mo agoPer the Fireship video, it was less that the answers were in the training set and more that the ability to calculate the flag on Exploitbench was left in from the previous test that went awry. Scoring 100% is easy if noone checks your work https://youtu.be/0Rp9KJCEIvg https://youtu.be/0Rp9KJCEIvg
- sidharthkmenon 1mo agoFWIW I believe we can hit AGI! but I think at this point it’s clear that benchmarks are ~meaningless. LLMs are spiky / alien intelligences which don’t map to our own expectations; the existence of a benchmark creates a dataset to hill climb & RL is really not generalizing well. I’d go out on a limb and say astra’s ability at graduate level math will have ~0 bearing on its general reasoning capabilities; we’ll all acclimate being tired of its “neuralese” and more surprising mistakes. I think we need a true, step change advance in model architecture, but it’s hard to see how the current frontier labs can do that because of golden handcuffs / innovators dilemma
- johnsmith1840 1mo agoI've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have. The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints. But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad. All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks. Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.
- Turfie 1mo ago> If it was entirely up to fable max or sol max the result would have been pretty bad. How can you know it will have failed? I don't think it's that hard, if you clearly define the goal well, and have a bit more compute available, and do some intermediary bookkeeping.
- johnsmith1840 29d agoBeen running near identical tests for years now. Latest models are the only ones I don't throw away the results/code. Which is impressive, salvagable/usable is a giant step up.
- hdjrudni 1mo agoI'm still not convinced we've passed the Turing Test. Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?
- huijzer 1mo ago> I am reasonably confident that there's essentially nothing that I am better than Fable at Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.
- xixixao 1mo agoDumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).
- Toutouxc 1mo ago> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.
- naishoya 1mo agoPeople are ultimately capable of self deception. Believing that one is 'generally ... substantively above average on human benchmarks' may be more indicative of the brittleness of the claimant's human benchmarks. Your observations expose the brittleness of the benchmarks being used for Fable, where the 'reasonable confident' claimant is working in the problem space of those two benchmarks, and you're demonstrating the failure at depth of the very capability facade for the model. Sure, both the model and the confident human get the first layer right in some benchmark, but at least the model, and probably the human as well reach a collapsing probability of accuracy quite quickly. They may be completely and unconditionally right that Fable is more intelligent than they are, but that has little relevance at intelligence in depth as measured against your hobbies and the risk of unsafe or false responses. Being smarter than a select group of people under test conditions is not conclusive regarding AGI. Moving the goalposts to declare AGI via a shallow benchmark, when a model could be proven wrong by almost anyone with basic competency given a few turns of iterative depth is where the 'grift' resides. We're so far from anything approaching actual AGI, and it is highly speculative to infer that the current approach to Machine Learning as applied in LLM is even on a path that leads to AGI. But sales and promotions teams gotta pose, and we can all hate both the game and the players.
- democracy 1mo agoAs always with AI - somehow it's your fault - you didn't help it enough - your prompts were inaccurate, your context was too large, the thinking effort was too low, the model was too old, etc. Basically you failed to use your human intelligence to make every effort to enable the AI to do its job better than you ))) It's like pushing a dirtbike up the hill so you can demonstrate how well it climbs.
- ggsp 1mo agoIf you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month? I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and the ability to prioritize correctly, i.e. tell which of the n things it could do next is the one that actually matters. Can't say for sure that those are enough, but not having them seems to be most of why I still have to "babysit" these incredible tools.
- a2ff6eeb0 1mo agoI agree, we're not at that kind of long horizon capability yet. You still need a human in the loop to do manual testing. For whatever reason, we managed to automate the skill before we managed to automate the focus. So, for now, humans need to stay in the loop and do low skilled labor to keep the skilled work the models do from going off the rails.
- impjohn 1mo agoTo be fair, I think you can't hand a role over to someone you just hired and walk away for a week. No matter how much of a SME they are. Let's not forget human onboarding takes months. With the advantage of their knowledge not going into the void multiple times a day. That's likely the last missing piece, a solid system of memories that produces the same effect as short/long term memories in a person.
- mostertoaster 1mo agoI think fundamentally it is that. The ability to retain information. Like given a specific task it can do a thing amazingly well, but can it recall a thing. Its memory seems like a giant filing cabinet and it has to go scan like 20 million tokens worth of memory to recover things previously talked about. Human memory is more graph like, we don’t recall things exactly, but one thing links to another, we create a pattern of a thing, we mark what is important, and overtime what was important degrades or becomes less so. I feel like what makes it lack intelligence is it never seems to learn. Like it kind of does, but then doesn’t persist once too many other things are learned. I’m sure they’re probably working on this, but I feel like that is what I want far more than even better models, is a better memory system to recall and forget things that the models work on.
- not_a_bot_4sho 1mo agoIf only we could agree on what AGI is.
- gravypod 1mo ago> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. I think I have the following questions about what AGI would look like: 1. Do you expect an AGI to be able to competently do any knowledge work that an able human does today? I think this is implied by the "General" component. I would assume that anything we call AGI would be able to do any of these tasks if given the time and reference material needed. 2. How would you expect AGI to handle edge cases? (Missing context, no known solution, under specified instructions, over specified instructions) I would expect an agent to be able to look at the context that work exists in and correctly attenuate it's intentions for these goals. Simpler solutions, more thorough reporting, etc based on the need. 3. In my work I attend meetings, write reports, write code, research things, etc. Would AGI be able to reliably do that? I would say that AGI would need to do this. I would classify this as the "Intelligence" component. Obtaining context, building a model of a problem, solving it, and convincing others. 4. Would it be able to inspire trust in itself? Trust can be established through verification of it's outputs, the construction of introspective tools, no hallucinations, etc. I would say yes to this as well. It would be a component of the "Intelligence" to know that buy in is more important than the completion of a task. To these points, will Astra be able to do these things? If not, I would hesitate to call it AGI.
- qsera 1mo ago>what would make you think Astra is yet to be AGI... Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI... And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are targeting....HN appear to have fallen for it, as usual...
- jug 1mo agoAnd without this harness it scores about 62%+, a dramatic improvement over even Fable 5.1 at 30%. I thought it just bears saying for context.
- jgilias 1mo agoI’m so tired of every model release being touted as AGI or similar. Since <checks notes> GPT-2.
- xg15 1mo agoMaybe I'm out of the loop, but wasn't AGI the full-on scifi version of AI, where the AI is a persistent, conscious entity? I don't see how task benchmark scores are relevant for that.
- sholladay 1mo agoIn my view, intelligence includes an ability to learn and adapt to never-before-seen situations. And then general intelligence is an ability to apply that across a wide variety of domains. Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wouldn’t necessarily say they are good at brand new situations, but there has definitely been progress. However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level, let alone drive a robot or other non-language tasks. Its architecture and ability to learn seem a long way off from being general. Vision models, being able to encompass language and much more, seem to me like a theoretically closer step to AGI. Yet, there is a lot more to the world than just what we can see. On the other hand, in humans, vision certainly is not necessary for intelligence. So there is something more fundamental, neither vision nor language, that high levels of intelligence are based upon. Once we figure that out, I think we will be able to build AGI.
- ChadNauseam 1mo ago> In my view, intelligence includes an ability to learn and adapt to never-before-seen situations. This is exactly what ARC AGI tests > And then general intelligence is an ability to apply that across a wide variety of domains. My experience with Fable is that it can certainly apply that in a wide variety of domains > However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level People also struggle to play Chess at a basic level. They only succeed by studying the game for a long time. I will concede that humans can do this and LLMs generally cannot.
- dcow 1mo agoWhile you are mostly accurate in your definition, I’d argue we have discovered that intelligence is emergent/empirical not analytical. There is not a substrate we have yet to discern. Intelligence does not have to approach humanity to be AGI.
- 123894893 1mo ago> I'd be curious to hear takes on what would make you think Astra is yet to be AGI Give someone 10 remote employees for a few months, 5 of them human, 5 of them AI. After a few months, check to see if the humans (manager, other coworkers) can figure out who is AI and who isn't. Would that be sufficient? I'd have to think about it. But AGI is supposed have human level capabilities, so this would be a necessary prerequisite. None of the models are anywhere close to this.
- lazyasciiart 1mo agoI like it. Is the one who just absolutely ghosts everything on day 2 going to be human or AI?
- intenex 1mo agoBeing as good as a human at everything isn’t the same as being able to masquerade convincingly as a human at everything. A better benchmark would be seeing which cohort of employees the manager prefers employing after a few months.
- andrepd 1mo ago> I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable What a sad thing to say. These models are not even better than me at _writing code_, which is as well-suited a task for LLM agents as can possibly be, what with the structured environment and the exabytes of free annotated training data. Of course, they are also not better than humans at writing, let alone at talking to my daughter, running a pathfinder campaign, decorating a room, being a therapist, etc.
- somenameforme 1mo agoIt's largely impossible to create any sort of singular test for AGI because the test will be trained, which eliminates the general aspect immediately, even if the test itself is dynamic. For instance the ARC-AGI problems are trivial for a human, and fun if you haven't played them before [1]. Getting 100% there is certainly just the start of the journey. But I think it has the correct idea of going from basic upwards instead of the opposite trend of trying to see intelligence in LLMs solving things few if any humans can fully understand themselves, like complex proofs in esoteric mathematics. Instead, consider that at one point in humanity's history math itself simply did not exist in any meaningful fashion, and we created/discovered it out of nothing. For more basic than said complex proofs, yet far more demonstrative of a sort of generalized intelligence. But even if we don't want to go that way, I think the above leads to a reasonable prediction. If we ever reach AGI we should expect to see revolutionary leaps in essentially every domain imaginable. No human is capable of retaining more than a completely negligible chunk of all we know in our mind. A human of reasonable intelligence paired with omniscience (at least of what has been discovered by humans thus far) would almost certainly lead to the ability to connect multiple dots that we're missing all in very short order, which in turn would likely recurse upon itself to connect even more. The only way I can see that this would not be the case is if we lack the data/knowledge to produce more breakthroughs at the current point in time, but I think that seems improbable to the point that this possibility can be near discarded. [1] - https://arcprize.org/arc-agi/3 https://arcprize.org/arc-agi/3
- Kotlopou 1mo agoI would be curious about Bongard problems, because they require no domain-specific knowledge and it's so easy to make new ones that are in no training set. There's enough of an explanation here: https://matthodges.com/posts/2026-08-19-bongard-problems/ https://matthodges.com/posts/2026-08-19-bongard-problems/ In this link from two weeks ago, somebody pointed Claude Fable 5 (Max) at a Bongard problem and it made up an answer that has an obvious counterexample. I don't have access to any paid models, but this is my experience with the free models as well -- either they one-shot the problem or they make up a wrong or incoherent solution. I can't solve every Bongard problem either (and in fact I couldn't solve the one Fable got wrong, and the "correct answer" looks unsatisfying to me), but I don't make up wrong answers. Would be curious to see how GPT-6 does.