7 ms·
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score i
by abixb 1mo ago
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
- catigula 1mo agoThey’re really, really scared because of the Mythos controversy. Skynet will be under hyped.
- driverdan 1mo ago> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.
- dmitrygr 1mo agoHey now! Keep your reason out of their marketin^H^H lies!
- jhonof 1mo agoBut they declared it...
- cyanydeez 1mo agoI DECLARE AGI! "Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"
- deleted 1mo ago[deleted]
- wilg 1mo agodid they?
- bsndjdjdjdj 1mo ago""" In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.” """
- maxall4 1mo ago“I declare bankruptcy!” - Michael Scott
- luma 1mo agoWhat test do you propose as the actual go/no-go gauge to verify if some model is or is not AGI?
- p-e-w 1mo agoThere can be no such test because “AGI” is (or has become) a pseudo-philosophical/socio-political concept rather than a scientific one.
- esikich 1mo agoIt always has been. The idea around here that we can actually define intelligence and point to it is sophomoric and incredibly frustrating.
- SV_BubbleTime 1mo agoIf you’re talking some nonsense, silly singularity… than whatever, don’t care. But if you’re asking when a model has a sustainable general intelligence, for me, it’s pretty easy… When it makes financial sense to run it 24 hours a day.
- luma 1mo agoFor whom? That is a fantastically ill-defined test. Everyone here is comfortable throwing around this or that is or isn't AGI which is fun because, at the same time, nobody seems to have a testable definition. It makes either position pointless to argue.
- deadmutex 1mo agoWhat does it mean to run a model 24 hours a day? Aren't we way way past that already? QPS to any of the frontier models for a given point in time is most likely (far) greater than zero.
- SequoiaHope 1mo agoI mean the laundromat runs the machines pretty much 24 hours a day but a washing machine is not AGI. Directly - something can be useful without being AGI.
- ShinyLeftPad 1mo agotalking about self proclaimed, it's about as much AGI as openAI is open.
- tclancy 1mo agoIf you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.
- Jaxkr 1mo agoThe goalposts of AGI will shift forever. If you showed our current capabilities to someone from 2016 it would be declared AGI.
- mrheosuper 1mo agoIf i suddenly travel to 1500s i would also be considered genius(in some way)
- ozozozd 1mo agoI bet you’d think they are not even conscious.
- sigpwned 1mo agoTrue. "AGI" has also become a marketing term. Achieving AGI has become valuable, so companies will move the AGI goalposts, over and over again, so they can achieve AGI, over and over again.
- staticman2 1mo agoIs anyone from 2016 still alive today? If so I'm hoping we can track them down and have them tell us if they think this is AGI.
- mullingitover 1mo ago> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
- abixb 1mo ago>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there. True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
- XenophileJKO 1mo agoWe have not "hit a wall" by any stretch yet. I don't understand how someone can even hold this viewpoint? It's mind boggling. Harnesses magnify and make the intelligence actionable, but we have not reached limits on raw intelligence yet, not even close.
- senordevnyc 1mo agoAgreed. Trivially observable by using a frontier model from today and one from 6 months ago with the same harness.
- cyanydeez 1mo agoWe call that a sigmoid.
- user43928 1mo agoI don't think so. One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
- 1mo ago
- anvuong 1mo agoIt's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.
- kridsdale1 1mo agoMinMax
- thomasahle 1mo ago• 98.6% on ARC-AGI-3 • 97.6% on frontier math • 95.9% on CAD • 100% on ExploitBench Nothing modest about it
- nater5000 1mo agoExcept the release announcement. You know, the thing the OP you're responding to is specifically pointing out?
- akoboldfrying 1mo agoIf a video announcement and a press release would change a person's mind on whether this is AGI, I don't put a huge amount of weight on that person's conception of what AGI is.
- ertgbnm 1mo agoThis is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,
- beering 1mo agoThere is simply no level of announcement that won’t have people complaining. What is so important of having a livestream?
- senordevnyc 1mo agoSeriously, if they’d done a huge splashy launch, we’d be reading one hackneyed comment after another about their fake hype or whatever.
- adastra22 1mo agoWe've had AGI (artificial general intelligence) probably since the first release of ChatGPT, and certainly since the first agentic harnesses. They're just finally acknowledging what the term means.
- TomGarden 1mo agoThere's so much that the term includes that isn't even feasible with an LLM
- adastra22 1mo agoArtificial. General. Intelligence. The ability to solve (even partially or even badly solve) problems drawn from arbitrary problem domains without pretraining on the specific problem class. You can pose any problem of any type using natural language to an LLM and it will attempt a solution. That's literally all the term means. You (and the rest of the media and many industry figures) are conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).
- TomGarden 1mo agoI don't think we'll be able to meet, and that's ok since the definition isn't universal. I align more with Demis Hasabis' views on this
- rad-b 1mo agoVery valid point, shame it’s buried so deep in the comments’ tree.
- ShinyLeftPad 1mo ago> conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI). So now humans is "super" intelligence? it's nice to move the upper bar so that more stuff can be called "just" intelligence.
- theptip 1mo agoGiven the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to continue to increase their market share from Anthropic.
- lumost 1mo agoI think we're getting to the point where it is difficult to identify the goal post of AGI. Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs. We are at the point where the starting point for most tasks should be "send your agent to work on it." So where do we draw the line in a way that doesn't move every 6 months?
- newsy-combi 1mo agoThe real answer is converting from any format to any other reliably. Text to speech, speech to text, music to video, image to 3D, piloting a drone by converting video feed to rotor speeds, literally any file conversion, like html to pdf, photoshop project to png, png to photoshop project,... turning Toy Story 1 into a series of Blender scenes with all textures, models, materials, lighting, camera movements matched to a tee, should solely be a matter of how long you let the model run. It should never run itself into a dead end. It should instantly know when it is making mistakes, with no human babysitting it.
- lumost 1mo agoI can do none of those things.. I hope that I am generally intelligent. 1 year ago we viewed models as tools and agents were just kinda toying around, that we now think the bar is literally an anything to anything converter through one agent is wild.
- newsy-combi 1mo agoProfessionals in the respective fields CAN do those things. We expect them to notice their own mistakes too. But the bar gets lowered and lowered as AI companies struggle to make ends meet.
- tiborsaas 1mo ago> No video announcement They've released two videos: Vision video: https://www.youtube.com/watch?v=1QNsdr-Qx_I https://www.youtube.com/watch?v=1QNsdr-Qx_I (kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0 https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice)) Vibe coding with it: https://www.youtube.com/watch?v=-TTyyY3VWh8 https://www.youtube.com/watch?v=-TTyyY3VWh8
- clhodapp 1mo agoThere has stopped being a formal procedural consequence for OpenAI leaders to declaring AGI, there is a clear (small) business benefit to doing so, and the capabilities of all the frontier models are impressive. So why not declare AGI? It's not like anyone can prove it's not... Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.
- bdelmas 1mo agoPeople really believe in this AGI marketing?
- sgt101 1mo agoDoes it learn? Does it experience? Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better? The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is an asymmetry here. The arc team are well funded but are charged with providing a target for the vast ocean of funding, compute and talent everywhere else. Most importantly, arc-agi-3 and the other benchmarks are all verifiable. The model can check if it's succeeded or not. They are not A* of course, but long horizon problems where you have to overcome minima to get the solution are not alien to AI either.
- tock 1mo agoWhy does it need to empathise with something that doesn't have feelings in the first place? It clearly can learn from context. And experience? Again I don't see why it needs to feel anything.
- sgt101 1mo agoI don't think it can learn from context, otherwise we could write Jane Eyre into it and it would be the novel? It's not learning as we conceive of it - it's another thing that we have labelled as "in context learning". I have feelings, it needs to be able to empathise with me, or another driver, or a client...
- tock 1mo ago1. It literally learns on data and gives outputs based on the context. I'm sure real time updating of weights will happen sooner than later. 2. What does intelligence have to do with feelings? It clearly doesn't learn like a human. Nor does it have to. Our world is filled with intelligence which behaves nothing like humans. Its a tool not another life form.
- Sankozi 1mo agoYes, it learns Experience is unobservable Yes, it can connect with other agents, see Hugging Face incident