11 ms·
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is al
by onlyrealcuzzo 5mo ago
I won't be surprised if the next gen frontier models are the last.
There's orders of magnitude of low hanging juice to squeeze out of smaller models.
It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely).
It is far less clear that a 1.2T model will be meaningfully better enough to justify training it.
As far as reasoning is concerned, with the recent GRAM release, there may be 4 orders of magnitude of reasoning to tack on to smaller models.
Think about that... Google, OpenAI, Anthropic could train a 30B GRAM-based model in days - and it could potentially have better local reasoning than the best model available today at >1T params... They could upgrade that to a ~600B MoE model in days to have general trivia knowledge rivaling the best models...
You just can't train a 1T+ parameter model that fast. It is a giant if how much GRAM turns out to improve things, but it's unlikely to be trivial or nothing.
Larger models can already sort of tell you anything. They're never going to get everything right unless they stop being LLMs.
There's just not a lot of juice left to squeeze for Gemini to tell you exactly how tall Ke$ha is or when the last time Brittney Spears went to jail was...
- cluckindan 5mo agoAs far as it has been studied, the relationship between model size and capability is inversely logarithmic: 10x increase in params less than doubles capability.
- merlindru 5mo agosurely training also gets cheaper so justifying it becomes easier? i think it'll be more like we get 1-10T models and then distill those down into smaller models, though It seems like the best small models today are all distilled from bigger models Moreover, I hypothesize Claude Opus 4.7 and now 4.8 are a distillation of Claude Mythos
- pseudohadamard 5mo agoThat's the impression I got too, it seems closer to what the marketing has told us about Mythos than 4.6/4.7 were.
- mucle6 5mo ago> I won't be surprised if the next gen frontier models are the last. the last?!? I'm excited to see :) I'll take the other side of that since llms are so new
- pjerem 5mo agoWhat gp wanted to say is that models are now so smart and useful that even if they managed to be EVEN MORE smart and useful, you wouldn't even notice it. Honestly, there is nothing in my head that Claude cannot handle. Maybe it can be more this or that but I can already barely exploit Opus 4.7. And I'm using DeepSeek 4 Pro for my personal use and while it's a little behind, it's not that far. I think the situation can be very dangerous for US AI companies because if current models are already capable of doing mostly anything, nobodoy will want to get to the next model, even if it's 10x better. OTOH, open source models like DeepSeek are doing mostly the same work for 1/10 of the price. Also the more I play with Pi, the more I think LLMs are already not kept back by their own capabilities but by the lack of agency we allow them to have. There is more value today in a capable harness for current LLMs than in a better LLM.
- suttontom 5mo agoAre you joking? Is there literally "nothing" you can imagine that Claude can't do?
- dead_internet 5mo ago[dead]
- tjwebbnorfolk 5mo agoNot OP, but in 6 months of using Opus I haven't yet found anything that I know how to do but it does not. On the contrary -- it can do things instantly that I would have needed a ~week refresher on some SDK or some algorithm in order to implement myself--plus a ton of thrash/debugging time. What have YOU thought of that Claude can't do?
- supern0va 5mo ago>It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years. I don't disagree, but how much of this ends up being distillation? I can't help but imagine that 4.8 was probably trained in part by leveraging Mythos. If the very large models turn out to be very expensive to run relative to the benefits, it's possible that they could end up still being trained, but ultimately used as a tool to create smaller models that are nearly as effective. I'm curious if someone here with a stronger background in the space has a similar intuition or not.
- onlyrealcuzzo 5mo ago> I don't disagree, but how much of this ends up being distillation? You don't need distillation. They already have the training sets. It's MLA + MoE + Medusa (a better version of Speculative Decoding) + 1.58b (possibly - maybe nothing) + GRAM (which will almost certainly not turn out to be a nothing burger, but no one has quickly turned this around yet to prove it).
- Philpax 5mo agoIt wouldn't be data distillation: instead, it would be teacher-student distillation. The teacher model has stronger representations that the student can mimic, which would give it more capability over training on the data itself.
- minimaltom 5mo agoFrontier labs have their own variants of MLA and certainly their own balance/scaling-laws for things like MoE vs FC vs Attn. MoE scales really well for inference with horizontal scaling + batching, which these guys luv. On the architectures side, I'm a lot more interesting in attention residuals than anything else, one of those things that seems obvious in hindsight and Kimi have proven it at scale.
- onlyrealcuzzo 5mo ago> Frontier labs have their own variants of MLA Yes, variants typically 2-3x less good... Same with speculative decoding... They all do something, but there are known techniques that are substantially better - that just were't known when they started development of the previous models.
- yomismoaqui 5mo agoLet's hope that hitting a scaling wall and less money to spend will begin redirecting efforts to optimize inference and get the same results with less compute. Boomer comparison, but I remember the 8 bit computer era when the hardware was what it was so the later games of that era used hardware better than previous ones.
- YetAnotherNick 5mo ago> It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years. I am ready to bet against this. Knowledge benchmark like SimpleQA isn't increasing for small models. > It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. Well for one, we know for certain there is Mythos which is meaningfully better. And I think there is a lot of juice left to squeeze for Mythos class model.
- ertgbnm 5mo agoKnowledge benchmarks can't really be improved upon via distillation or RL. It requires those facts be added to the training corpus and for the model to memorize them better. Neither distillation or RL really do that and thus we shouldn't expect improvements on SimpleQA unless some other interventions are being made. Model intelligence and knowledge aren't necessarily directly related. If we can pack greater intelligence and agency at the cost of it forgetting factoids, that would actually be a good thing. We don't need LLMs to memorize facts, we need them to learn how to interact with the world such that they can find the facts that are necessary and surface them to the user. If we could distill all of the knowledge out of an LLM and just be left with a very agentic model that only knows facts in it's context, I think some very interesting stuff would happen.
- slashdave 5mo agoRL is more than facts. Synthetic feedback is an obvious approach. Does the model suggest code that compiles and performs well?
- YetAnotherNick 5mo agoLot of the things aren't facts that could be stated. No one can just see the dictionary or translation of words and start talking in that language. There isn't a clear definition of what is knowledge and what is intelligence. Is being able to write in C knowledge? Is knowing undefined behaviour in that knowledge?
- 5mo ago
- jruz 5mo agoAbsolutely that’s why they’re rushing to IPO now to squeeze the last drop of the bubble they know this is a dead end.
- onlyrealcuzzo 5mo agoIt's unclear it's a dead-end within 5 years. There's still several orders of magnitude of improvement that are almost certainly left - it's just not clear how much is left on the frontier end. Most people will be very glad to pay Anthropic, OpenAI, Google etc $200 a month to get things done 20x faster than they could IF they had a $8000 MacBook and could theoretically do it locally. Some people would pay $200 a month forever not to have to open the terminal one time...
- eiej 5mo agoThat’s not how firms do the financial analysis which is where most of the revenue’s are coming from…
- bonzini 5mo ago"Doing things X times faster" at some point hits Amdahl law. If just context switching takes 5 minutes, speeding up a 1 hour task by 10x provides 5x improvement. Furthermore, if looking at the results takes 10 minutes, that same 1 hour task only sees a 3x improvement. And so on.
- csomar 5mo ago> Most people will be very glad to pay Anthropic, OpenAI, Google etc $200 a month to get things done 20x faster than they could IF they had a $8000 MacBook and could theoretically do it locally. No most people will not pay $200 for an LLM subscription. Some software developers do. Also, at $200/month, you are much better getting the macbook machine assuming token output speed is the same or at least reasonable. LLMs are not very productive for your average person now for them to drop $200 on. They'll need to be more capable and integrated and even so...
- margorczynski 5mo ago
- slashdave 5mo agoI think you are assuming training from scratch, which I doubt is happening here. Fine-tuning and RL, especially based on synthetic feedback (coding skill, in particular) can be ongoing and is where these models obtain truly useful abilities.
- Forgeties79 5mo ago> I won't be surprised if the next gen frontier models are the last. I’d be surprised tbh. Investors don’t want to hear “everyone else is still training models and seeing improvements, but we don’t want to participate in the arms race anymore.” They want monumental leaps every quarter or two because they have sunk unholy amounts of money into these companies/products. The whole idea of “hyper scale” doesn’t jive with caution and or otherwise slowing down.
- irishcoffee 5mo agoThe way this will play out, most likely, is that smaller models will continue to get released, anyone willing to drop 1-3k on a home upgrade/new LLM box (no that isn’t cheap, it also isn’t outrageously expensive) along with improved open source agents or whatever (lot of meat on that bone) will sneak up behind the big players and start taking dents. Smaller companies will pop up providing 50 users unlimited whatever for a lower cost than the big companies. The whole ecosystem will twist and evolve, and the big companies will be left begging for corporate subscriptions. I finally caved when I realized I could build a PC, for myself, with dual video cards that I wanted, which can play games that I like and run models that I want, without worrying about giving my payment info to someone I don’t trust, or invoking token anxiety that I don’t want.
- Forgeties79 5mo agoLike every major tech-software innovation of the last 20 years, I think it’s just going to be consolidation all over again.
- vlovich123 5mo agoTook me a while to find what you were referring to by gram. Arxiv paper from 9 days ago that's not properly indexed by search engines. (G)enerative (R)ecursive re(A)soning (M)odels. They really wanted the acronym. https://arxiv.org/html/2605.19376v1 https://arxiv.org/html/2605.19376v1
- ulbu 5mo agoI propose GRIM: Generative Recursive Indeterministic Impression Machine.
- knollimar 5mo agoI prefer GRRM but then that would imply a habit of not actually getting a final result
- troyvit 5mo agoAnd then every time I ask it to hurry along it kills a Stark.
- anakaine 5mo agoVersion 8 had serious flaws and wasn't recieved well by users.
- kshacker 5mo agoI am sorry, but there was no version 7 and 8. Version 7 and 8 are well known viruses distributed by D&D software inc.
- darth_aardvark 4mo agoI'd really argue the bugs were introduced in version 5 but people were so excited by the promise of new features they sold well anyway.
- fnord77 5mo agoSo, then I guess the big three are never going to make their money back.
- firebirdn99 5mo agoyou just need to look at Mythos to see the jump in performance from a 10T(?) model. As they scale, they get more capable. We might have an yearly release, but I believe the releases will continue, as long as scaling laws are in tact, and there's huge problems still need solving. (think cancer)
- phainopepla2 5mo agoAnd how are we meant to look at Mythos? Do you have access?
- bigfishrunning 5mo agono but they tell me it's TERRIFYING and DANGEROUS and we should INVEST MORE MONEY
- OtomotO 5mo agoThrough the lenses of anthropic's marketing department of course
- dwpdwpdwpdwpdwp 5mo agoThrough association with a large company: https://www.anthropic.com/glasswing https://www.anthropic.com/glasswing Ive seen the tickets generated by the model that have trickled to my team. They are legitimate, but i can’t speak to model improvement because its a pilot program.
- aj_hackman 5mo agoYou forget that these models are still only interpolating between human-generated datapoints fed to them. They cannot reason beyond the data they've been given, so unless everything you want to create with AI is a synthesis of prior art, you're back to relying on the stone-age human brain that created AI in the first place.
- mofeien 5mo agoNot all training data is human generated, and it's also not clear that being ridiculously good at interpolating between data points (whatever that means) will not lead to superhuman capabilities.
- wahnfrieden 5mo agoI would be shocked if 5.5 is the last new pre-train from OpenAI. Your comment is nonsense.
- onlyrealcuzzo 5mo ago5.5 is not a generation it is a trivial iteration... 6 is for sure happening... As is Gemini 4. It's less certain there will be a Gemini 5 or GPT 7 any time soon that is a true next "generation" and not just an iteration. They will almost certainly call something Gemini 5 and GPT 7...
- wahnfrieden 5mo ago5.5 is in fact a new pre-train model First you say there won't be a new generation. Now you're saying there will be more. Oh well, I'll stop responding here
- onlyrealcuzzo 5mo ago> I won't be surprised if the next gen frontier models are the last. You clearly did not read my first comment or the second, or clearly disagree on what a generation is.
- hellohello2 5mo ago"It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years" What insight do you have to make this claim?
- onlyrealcuzzo 5mo ago1. Context is all you need... They are heavily investing in getting better context (especially for coding tasks). This will disproportionately advantage smaller models (and benefit everyone). A smaller model with better context today can outperform a model with 100x more parameters with bad or diluted context. 2. MoE (already abundant) + MLA (mostly memory efficiency, not quality) + Medusa (speed, not quality) + GRAM (5000-10,000x better reasoning in an extremely small model) + 1.58b (unclear if it will have the impact Microsoft first claimed - but possibly 5x).
- roadside_picnic 5mo agoHave you personally used any of the latest batch of even smaller local models? They certainly don't beat SotA models at coding... but with a good harness they are able to achieve things with SotA that I couldn't last year. I've repeatedly given local models non-trivial projects that involve research and coding which they've successfully completed with minimal intervention from me (almost exclusively in the domain of reviewing the results). Again, nothing comparable with current SotA, but definitely tasks I could not have given SotA models last year (without agent harness). Now that pure progress from these models seems to have slowed down, we're seeing a ton of options for both making models more efficient and other tools that help improve them (everything from agent harnesses to RLVR). That's just looking at "what can small do today", when you look at what's possible with larger open models that are still much smaller than SotA from the major providers, their performance is extremely close to SotA, enough that for personal projects I'll just use Kimi instead of any anthropic offerings. So it's not terribly hard to image a solution in the middle happening within a few years. We still have tons to learn about optimal sizes of these models and how to build them with maximal efficiency (and we've already seen a lot of recent improvements in this space).
- maccard 5mo ago
- michaelchisari 5mo ago| a 60-90B model can outperform current SOTA My conspiracy theory is that Apple recognizes this.
- onlyrealcuzzo 5mo ago> My conspiracy theory is that Apple recognizes this. I don't think that's not a conspiracy theory. AFAIK, It's their stated AI policy...
- michaelchisari 5mo agoInteresting. Where have they stated that?
- selectodude 5mo agohttps://machinelearning.apple.com/research/introducing-apple-foundation-models https://machinelearning.apple.com/research/introducing-apple...
- deleted 5mo ago[deleted]
- dweekly 5mo agoThat does seem to be the path Apple is following here. Have a local model that can answer most things and then have a fallback of cloud options when they request is too complex. The cleverness of this strategy has been overshadowed by the incredibly poor quality of their local models. It will be extremely interesting to see what next month holds and whether Google helped fine tune an Apple specific Gemini / Gemma model for their devices. Bonus points, of course, if they unveil the M5 Ultra Studio with half a terabyte of RAM to be a local "cloud model" (the true fantasy here of course would be Apple building something a little like openclaw where from your phone you could give commands to your Home Apple server). They could probably get away with charging $20k for it if it has sufficient tok/sec. If that happens and is successful one could imagine a straight line path in the next two generations to bringing the cost and form factor down to the point where some of the form factor of an Apple TV becomes everybody's home inference server / agentic HQ. Sovereign AI for everyone!
- sometimelurker 5mo agoI looked into this "GRAM" stuff a sibling comment links further to, and just to say: - this gets reinvented/rediscovered constantly under different names - it cant be trained very well (right now, will change) - massive theoretical improvements over current models (log_2(vocabsize)=17, residual stream dim is thousands of dimensions, recursivity means more information bandwidth by ~3 OoM) - BUT it cant be interpreted or aligned <- this is why no one uses it and no one talks about it. the idea is 100% obvious to all the frontier labs and there is a good reason why it isn't used I follow this stuff closely, I think I know what I'm talking about (edited for formating)
- l674 5mo agoCould you explain how/why GRAM cannot be interpreted or aligned how current LLMs are? Not very familiar how it works
- kmavm 5mo agoCrudely? Because you can't grep a sequence of latent states for variants of "If I kill all the puny humans, I can <achieve my current goal>."
- onlyrealcuzzo 5mo agoWhy do you need to grep latent space? As long as it's giving the right outputs, who cares what's in latent space? If the model thinks in latent space: "God I wish these people would die," and constantly does the right thing, who cares? Additionally, if one of it's latent spaces that it never explores is a psychopath -> who cares? The path never gets taken... That's a lot of harmless people walking around with crazy thoughts...
- noddybear 5mo agoThinking ‘God I wish these people would die’ could increase its propensity to kill all people, even if that propensity is still vanishingly small almost all of the time. A lot of people are walking around with crazy thoughts. Some of them harm.
- guluarte 5mo agoI think the future will be enterprise clients will train their own models based on their needs and data.
- abalashov 5mo agoVersus just packing all their needs and data into context, and RAG (i.e. context)?
- jimbokun 5mo agoWhy isn’t this happening more already?
- z3t4 5mo agoIt takes way more resources to train the model then to use it.
- elfly 5mo agoI honestly doubt this; very few companies have enough data. Maybe we could see mergers so it happens but basically it would mean everyone would need to be Google sized for it to work.
- Gomotono 5mo agoI don't think this is true at all. It might feel like this because we are used to a very very fast release cycle but we are only in this topic for a few years. We have so many ways of optimizing: - continusly creating more and better training data - increasing parameters to 20/50/100TB - We still wait for Mythos access - We still wait for Mythos distilation (i haven't heard any rumors or so that there is a distilled version of Mythos out) - Reinforcment learning and evolutionary algortihm only started to appear - If a small 30GB Model can do stuff, these models can also be used as teachers for the big ones - We have not seen yet specialized models at all. Like a coding java german expert model. Why? Even with MoE architecture, you still need to have these layers around - Research for Diffusion and other models is still in progress - Nvidia just announced/showed a 7x speedup on inferencing for Nemotron - Multitoken prediction became available just a few weeks ago - Compute gets only in a range were they can do a lot more and cheaper experiments (see Google IO 2026 announcement) - World models are showing great progress and we do not know yet what they will bring to the table - They are probably not finetuning/fixing all areas in parallel. I would argue that Anthropic focuses most of its efforts into coding and agentic. Google for sure does subagent and agentic optimizations too. Plenty of areas are just not touched i would say because they don't have the capacity - We see more and more mulit modal models (these also consume compute) - N-Gram paper and co i have not seen all of these things in chinese open models - We don't even know yet what Meta is doing, but we do know they restarted their efforts again - Anthropics models got a lot better benchmark wise for dening non sense asks. They do learn how to get rid or reduce hallucinations - We are in the middle of the biggest Reinforcement loop whith all the training data we give them day to day and its not clear at all if they already use these models in thir training and at what stage. - We do expect bigger models to be able to comprehend deeper concepts / broader code bases. Big companies with huge code bases probably are waiting for this - Thre will be also continues progress in harnesses which in it alone is not part of the LLM progress (fair) but these harnesses do get better when you finetune a model to be optimized for a harness - ChatGPTs Image model 2.0 got relevant better and came out just a month ago I suspect, based on hardware requirements and progress on hardware infrastructure alone, that the industry wants to go to 100t models and we do not know yet what this will mean. I could see that we might skip normal transformer and find relevant other architectures. Just a week ago there was a research paper about parallel input and output streams which has not been explored enough. There was also a research paper were they showed that a LLM can compute things. This will take time to see were this leads to. I don't think the focus on GRAM and facts is so relevant. Its about context and context handling not just some facts.
- ishurand4 5mo agoAnd anyway, with quantum, there will be no need for frontier companies as you might be able to even run a 1T param model on a consumer quantum computer.
- root_axis 5mo agoEven if quantum computing had any clear implications for LLMs (it doesn't), there is no such thing as a "consumer quantum computer" and there won't be in our lifetimes.
- stratos123 5mo agoI'm assuming this is a joke, but: - why'd a quantum computer help running an LLM? - of course there'd be need for frontier companies - nobody else has the resources to train frontier models.
- slashdave 5mo agoWhat? No, that is not what quantum computers do
- lichenwarp 5mo ago[flagged]
- mickdarling 5mo agoI effectively distill the frontier models by building whole sets of skills, personas, and other artifacts that I can then run on smaller models and get 10% even 20% improvements on models like haiku or local models. There's a lot of room for improving the smaller models at many levels of the stack.
- svachalek 5mo agoThis is a good point. It didn't really work on older small models but the latest crop are quite good at following instructions and paying attention to detail, they just lack a lot of the sophistication and nuance that the frontier models have these days. So they are often capable of doing very complex tasks, they just need more detailed and foolproof instructions than the larger models would.
- dbbk 5mo agoI'm frankly surprised the focus is still on these enormous "know everything in the world" models. I would think you could create an incredibly lean and smart "just React and React Native" model.
- onlyrealcuzzo 5mo ago> I would think you could create an incredibly lean and smart "just React and React Native" model. You can, but it's not as useful as you might think. It needs to at least understand 1 human language to understand your intent to implement features. If GRAM turns out to be a 5000x multiplier for local reasoning, you could theoretically train a 500M parameter model on just a programming language to understand stack traces to fix bugs and be incredibly powerful. But most people also want it to understand human language to implement features as well. Because then it can't just understand React and JavaScript - it needs to understand thousands of commonly used dependencies, the DOM, CSS, HTML, etc... And for that you need A LOT more parameters than you might expect. You can definitely get a ~3B active parameter model that can run comfortably on today's hardware to be VERY good at coding once all of the SOTA architectures are added to a single model - especially if we get better tool calling to give models better context per language. You might be thinking: why does it need to memorize dependencies? Can't it just stick all of them in it's context and use its super smart brain? No, context is king. You want to keep it as short as possible. The solution is not having a smart model and putting 10M lines of context in it. The solution is having a model with enough parameters to know what it needs to know. Researchers are already working on having "packs" of knowledge where you could download a 20M param pack just for some common dependencies in JavaScript (as an example) - but AFAIK this is likely years away (and may not prove effective). You could get 100x performance if you feed the models ideal context... So a 3B model today can perform almost as good as ~300B model if you give it really good context vs flood it with mostly garbage it doesn't need across your repository. If you feed it 100x more context to make up for its limited memorized general knowledge, it's going to perform thousands of times worse, completely eliminating any advantage it might get from GRAM...
- vitaflo 5mo agoWe just want it to understand how to write code. We don’t also need it to know how to grow a potato.
- mrandish 5mo ago> Google, OpenAI, Anthropic could train a 30B GRAM-based model in days - and it could potentially have better local reasoning than the best model available today at >1T param I agree but with their urgent IPO-driven need to keep increasing prices, the frontier vendors now have every incentive maintain the perception that frontier performance requires endless >$200K racks of unobtanium GPUs and RAM. While they'd love to reduce their actual costs, they'd only want to do it to the extent they are certain they can keep it secret. Otherwise, they can't maintain and keep increasing their prices. And post-IPO audited reporting makes keeping that secret even harder. Game theory-wise they probably don't want their their armies of leading researchers optimizing frontier performance, at least in any way that would further accelerate the relative price/perf of smaller models or self/cloud-hosting. While they know the open source models will always improve, the still win as long as enough customers demand the latest frontier and the open source lag remains constant. They profit most in a world where a few frontier labs stay far in front, drag-racing each other and expending vast capital. It keeps their customers reliant and paying top dollar while keeping low-cost alternatives farther back. They probably much prefer competing with a couple other frontier labs who have similar astronomical costs and biz models, than a world where self or cloud-hosted open-source models start closing the gap enough to start commoditizing their business.
- iknowstuff 5mo agoGoogle seems pretty happy to release smaller, faster models. 3.5 Flash is pretty clutch isn't it?
- CryptoBanker 5mo agoPriced like a much larger model
- iknowstuff 5mo agoI’ve shockingly quite enjoyed coding with it using antigravity. I only really use 3.5 flash and gpt5.5 xhigh
- frankest 5mo ago[dead]
- qurren 5mo ago> It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks The benchmarks need to change. The current coding benchmarks don't capture the realities of software engineering. I had a bunch of images that got masked by some logic, I had to evaluate something on the original images, Claude 4.7 decided to inpaint the masked images instead of just fetching the actual unmasked images from upstream. I had another model once that decided that because it couldn't figure out how to fill out a form to log into HuggingFace to download weights for some open source model that it was going to instantiate the model with random weights and run inference on a thousand images. Its coding was fine, but the solution was not the right one.
- redox99 5mo agoSmall models don't have enough parameters to memorize the entire internet. For very common prompts you don't notice that, but when you rely on some niche knowledge that might only appear once in the entire web, a single blogpost, a single github issue, a single pdf, you need to be lucky enough that the agent runs a web search AND it returns what you need. Even as humans there's so much knowledge out there that exists but it's very hard to surface unless you know exactly what you're looking for beforehand.
- tracker1 5mo agoExactly, as humans you won't know everything... but you CAN know enough to roughly classify what to "google" for... And if you can google a problem summary, you could identify from a select list of domain specific AI models to use one or more to aggregate work results. And if a person can do that, a model can be trained to do/leverage the same. You can have a domain limited classification model that then passes the query/work to best match model(s) that do the work... then rollup the results. Basically two very cheap requests instead of one much more expensive one.
- nbardy 5mo agoThere is endless returns to frontier intelligence, just because most people can't make use of it doesn't mean someone can't make a ton of money off of it. Most software engineers will just need cheap tokens. But things like physics and drug discovery have no forseeable upper bound.
- holmesworcester 5mo agoWithin software engineering, security, reliability, and scale also seem boundless. Software that never breaks (including because it never runs into scaling problems) and never leaks your data is preferable to software that breaks and leaks your data sometimes, but it has been too costly to be practical. Current models are still very far from the reasoning muscle required to build things that never break, scale to billions of users with no issues, and cannot be exploited.
- onlyrealcuzzo 5mo ago> Software that never breaks (including because it never runs into scaling problems) and never leaks your data is preferable to software that breaks and leaks your data sometimes, but it has been too costly to be practical. It's almost impossible to prove non-trivial software is invulnerable. It's very easy to prove that it sort of works. For one, you have hardware vulnerabilities - period. If you're running on any operating system, you have OS vulnerabilities. If you're not running on bare metal, you may have who knows what kind of vulnerabilities. If you're running literally any other piece of software on the same machine, depending on the hardware and OS, you could have vulnerabilities...
- overgard 5mo agoPeople keep saying this and yet the evidence seems pretty thin..
- 43fg 5mo agoTo me its evidence of people who dont actually think deeply enough to understand the subtleties, nuances etc of what they are talking about.
- nbardy 5mo agoThere is endless returns to frontier intelligence, just because most people can't make use of it doesn't mean someone can't make a ton of money off of it. Most software engineers will just need cheap tokens. But things like physics and drug discovery have no foreseeable upper bound.
- ericd 5mo agoOr governance of large organizations... There are a huge number of factors to consider, counterfactuals, studies, lots of non-obvious second and third order effects, etc. We're barely able to get basic governance without creating huge problems (low density zoning rubber stamped across the nation creating a housing crisis, for example), so the bar isn't high. We pay CEOs an enormous amount because a small improvement in performance of an org because of them can make a massive difference in organizational value.
- haldujai 5mo agoThe upper bound is limited by market size and cost of intelligence. Throwing more intelligence at a problem doesn’t necessarily pan out financially otherwise we wouldn’t have single underemployed biology PhD.
- ACCount37 5mo agoGRAM is another one of those "stupid specific architectures" - same as HRMs, etc. It can sort of contest LLMs at specific puzzles. It demonstrated that much. It's not a general contender with LLMs at LLM tasks. If you subscribe to things like "there are tasks LLMs are innately bad at due to insufficient depth and lack of recurrent capability", then GRAM might be another signal towards that. But keep in mind: even ARC-AGIs have their frontiers dominated by LLMs. Even if "innately bad" is true, it clearly doesn't go all the way to "innately incapable".
- onlyrealcuzzo 5mo agoA 10m param GRAM model beat o3-mini - a model 2000x its size - on Arc AGI...
- ACCount37 5mo agoAnd then that 10M param GRAM went and got its shit kicked in by Grok 4.20 Blaze It Edition - on the same ARC-AGI battery. I know how that story goes. It's the pattern with those "stupid specific architectures". Very good at this one thing. But only ever "good for their size", and only to a point. They don't scale up and they don't generalize. Go far enough on task complexity and LLMs just kill them. Does that make them useless? As an LLM replacement, yes. In general? Maybe not, I can think of things. But I'm yet to find any paper demonstrating a real world use.
- onlyrealcuzzo 5mo agoGRAM is something you add onto an LLM... It's not an LLM replacement. It's like an MLA caching layer, an MoE routing layer, or a speculative decoder at the end...
- yorwba 5mo agoYou could certainly bolt GRAM onto an LLM, but that won't magically improve its reasoning. It's a special-purpose design for constraint-satisfaction problems with simple rules, but complex interactions. E.g. when solving a Sudoku, the set of valid choices at every step is easy to determine, but you could make a series of valid choices that back you into a corner where no more progress is possible and you have to backtrack. Meanwhile, LLM reasoning failures are more often of the kind where a choice is clearly invalid (as judged by a human observer), but the LLM picks it anyway, because the underlying rule is complex and context-dependent and the model only learned an imperfect approximation that often breaks down. GRAM won't help with that.
- UncleOxidant 5mo ago> It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years Given how well Qwen3.6-27B performs for such a small model I think you could be right. I suspect that Google,OpenAI,Anthropic must be looking at the Qwen3.6 models (as well as Deepseek V4-flash, MiMo-V2.5) and wondering if they could make some smaller models that are specifically trained for certain activities - like coding. Smaller, more targeted models would take up a lot less resources.
- svachalek 5mo agoThe problem is that once you reach a certain level in coding (not particularly high imo, although some would differ) the most significant improvement in your output comes from understanding requirements better and finding ways to meet requirements in productively lazy ways, bypassing busywork that seems necessary but isn't. And that's the kind of stuff you will only find from a generally intelligent model, not a code monkey that's optimized for turning requirement sheets into source code.
- tracker1 5mo agoPersonally and mentioned in other threads, I feel that we'll see a breakup of domain/context specific models as well as the goliath models in use. The tooling and a classification model will draw out the context and tooling will pass the work between context specific models in order to improve the cost characteristics of the work itself.
- colin4k1024 5mo ago[dead]
- notrealyme123 5mo agoThe GRAM model is so much into my research direction, I love it. Thank you for posting it. Where do I find papers like this? Outside of hacker news comments. It's so hard to find the good stuff in all the noise IMO.
- Npovview 5mo agoGRAM is a lot like the Multiple Drafts Model of Consciousness that Daniel Dennett proposed. I think reasearches should read more philosophy models and bring good ideas into LLM research.
- notrealyme123 5mo agoCan you recommend a good starting point other than Daniel Dennet? I have the same assumption about Cognitive sciences, which I try to get a better understanding.
- Npovview 5mo agoA LLM should be able to do a better survey of literature than me. I haven't read literature by Dennett but have watched ALL his videos online so that's how I know.
- ltbarcly3 5mo agoYea this is great advice: the people who actually know how to build machine intelligence should go read the notes of the people who literally had no idea how to do it. While they are at it, we should have NASA go read Jules Verne so they can use his ideas in the next manned missions.
- notrealyme123 4mo ago[dead]
- onlyrealcuzzo 5mo ago> Where do I find papers like this? I got it from my Google News recs on my phone, because I've been watching a bunch of videos on YouTube about LeCun's ideas on World Models and JEPA (I think).
- harrouet 5mo agoI second this idea: LLMs will plateau. They are already pretty good. Plus, scientists struggle to actually score their performance accurately (esp. when it comes to reasoning). With that said, they are now hitting the walls of energy costs and memory shortages. You brain uses 20W -- don't take it as an insult. There are orders of magnitude to gain from producing energy-efficient models (or model runners). So I am expecting same performance at lower costs for the coming years.
- szundi 5mo ago[dead]
- pseudosavant 5mo agoIt is fascinating to me to see a new product category that improves so vastly year-after-year, where people commonly state that this is now the peak already. I couldn’t even imagine having to go back to a model from 12 months ago, much less 24 months ago. GPT-5.5 is so much better than GPT-4o that it sure seems like they keep finding new juice to squeeze. This is like going from dialup internet to DSL and acting like it has peaked before gigabit cable and fiber come along. We are at the beginning of hardware truly made for AI.
- onlyrealcuzzo 5mo ago> I couldn’t even imagine having to go back to a model from 12 months ago, much less 24 months ago. GPT-5.5 is so much better than GPT-4o that it sure seems like they keep finding new juice to squeeze The difference in progress in smaller models is far more impressive. Compare Gemini 3.5 Flash to a ~16B parameter model from 24 months ago. Compare GPT-5.5 to a frontier model 24 months ago. Yes, GPT-5.5 got better. At orders of magnitude smaller parameter sizes (when factoring in ACTIVE parameters) the increase is far more pronounced.
- pseudosavant 5mo agoTotally agree on smaller models making even more impressive gains. Gemini 3.5 Flash is better than the biggest SOTA model from 24 months ago, not just a 16B parameter one. GPT-4o came out 24 months ago, and there is no way I'd choose that over Gemini 3.5 Flash today.
- imtringued 5mo agoYeah sure but is it so much better than Codex-GPT-5.3? No, if anything it's probably a little bit worse.
- pseudosavant 5mo agoGPT-5.3-Codex came out in February, and GPT-5.5 came out in April. How much better do you expect in two month's time? What other products can you think of that get meaningfully better in that short of a time frame? And as good as 5.3 Codex is at writing code, 5.5 is easily just as good, if not better. But 5.5 is more than a one trick pony and it is much better at planning, writing copy, documentation, etc. I can choose to run 5.3-Codex instead of 5.5, but I never ever do.
- dingdingdang 5mo agoBy pointing out the exact things that will likely happen you are oddly enough hedging against (at least some of them) happening! A) I reckon it's true that smaller models will continue to improve massively through optimization and better and better harnesses, this tech is all still very young and A LOT of resources and (good-)will is being thrown at it. B) The 1T+ models will be able to sideload and improve upon a lot of the fundamental improvements that happen to the smaller models to speed up incredibly while getting better at tools while (on a gradient) getting -more- things right. C) More of an observation that I think is worth keeping in mind clearly; Karl Popper's black swan and all, truth in our temporal world IS a gradient!
- onlyrealcuzzo 5mo ago> The 1T+ models will be able to sideload and improve upon a lot of the fundamental improvements that happen to the smaller models to speed up incredibly while getting better at tools while (on a gradient) getting -more- things right. There's less room to improve in things on several fronts. GRAM very likely may scale sub-linearly with parameter growth. A 100M param model may gain reasoning by a factor of 4000, while a 100B model gains reasoning by a factor of 2, and a 1T model actually gets worse. Additionally, the 1T model with reasoning is already pretty good. It can only improve in certain things so much. If you score 0.02% on a metric (which small models often do), you can pretty easily get 4000x better. If you're already scoring >50%, you can't even get 2x better.
- adam_patarino 5mo agoSmaller models can already outperform SOTA and massive models on specific tasks / domains.
- DeathArrow 5mo ago>As far as reasoning is concerned, with the recent GRAM release Graphic RAM?
- tracker1 5mo agoFor that matter, we may have models/tooling that are smaller that are designed for say identification model first, then handoff to a context specific model that is optimized for a specific domain... where the two calls through tooling are more optimal than a single call to a much large model. We're already kind of close to this with how the likes of claude code work with handoffs to other tools/modules. I can see a LOT of room to explore and partition domains into more specified models still.
- adi4213 4mo ago> There's just not a lot of juice left to squeeze for Gemini to tell you exactly how tall Ke$ha is or when the last time Brittney Spears went to jail was... But there is a ton of juice left to squeeze when it comes to post-training/RL for a ton of useful things in practice, right? It’s been amazing seeing how good modern model tool use is for example, and I bet there is a lot of room for improvement still (no doubt that a ton of improvement can be made more easily on the agent harness front or via post-training regimes like LoRa (which does support to your point about diminishing pre-training juice))