3 ms·
> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. Hot take: These models are never going t
by mullingitover 23d ago
> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
- abixb 23d ago>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there. True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
- XenophileJKO 22d agoWe have not "hit a wall" by any stretch yet. I don't understand how someone can even hold this viewpoint? It's mind boggling. Harnesses magnify and make the intelligence actionable, but we have not reached limits on raw intelligence yet, not even close.
- senordevnyc 22d agoAgreed. Trivially observable by using a frontier model from today and one from 6 months ago with the same harness.
- cyanydeez 23d agoWe call that a sigmoid.
- user43928 23d agoI don't think so. One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
- abixb 23d agoI think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.
- NiloCK 22d agoIf someone could tune models of that size to have comparable effectiveness at much much lower costs, they would have done so by now. "The harness improvements are the real sauce" is like a sincere "It's gotta be the shoes" take about Micheal Jordan. (For the younger: that line was from a series of Nike ads where his skills were being explained)
- mullingitover 22d agoIt's more like we just invented ball bearings. We just jumped from standard to industrial grade, and precision grade is on the horizon. All kinds of new possibilities have opened up, cars can go a mile a minute on these things! Surely if we keep increasing the precision at this rate, we'll defeat friction once and for all.
- chrismarlow9 22d agoI don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
- frabcus 22d agoYes, the very explicit plan of both OpenAI and Anthropic is to use the not particularly efficient LLMs to automate their own AI engineering. That seems to be going well - on coding front and model tuning front so far. They have more planned. And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain. It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it. It's scary, TBH.
- m11a 22d agoI recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code. But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.
- seanw444 22d ago> but I didn't think it'd solve maths problems Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
- 22d ago