4 ms·
Indeed. It's pretty interesting to realize after implementing GPT-2 that the frontier models are scaled up versions of that, with various tweaks to improve perf
by jfim 4mo ago
Indeed. It's pretty interesting to realize after implementing GPT-2 that the frontier models are scaled up versions of that, with various tweaks to improve performance, model-wise.
The secret sauce though is all the datasets, RL training, knowledge of what works from doing all kinds of ablation experiments, and a massive compute moat.
- achrono 4mo agoHow do we know that today's frontier models are merely scaled up versions of that? Genuine question, since the labs have narrowed what they share over the years to now almost nothing, in terms of how the model was trained and how it works under the hood.
- ai_slop_hater 4mo agoNo they are clearly not just scaled up versions of gpt 2; there are different LLM architectures like mixture of experts etc that appeared relatively recently. I am not an expert though, far from it.
- otabdeveloper4 4mo agoMoE and such are basically performance enhancements, they don't make the model smarter.
- yababa_y 4mo agoseparately trained experts can surpass performance in their activated regime and DOES result in a smarter model, the Claude system cards talk about this and eg there is https://openreview.net/forum?id=iydmH9boLb https://openreview.net/forum?id=iydmH9boLb to read...
- jmalicki 4mo agoPerformance enhancements are huge though. If you can make the existing model faster, you can then save your inference budget to then make your model bigger, which then makes it smarter. A lot of how smart the models can be comes down to budget. If you can make your existing thing cheaper, you can instead make it bigger for the same price.
- TheHalfDeafChef 4mo agoNot really “smarter” though? It’s just a big probability engine. (Not trying to flame bait or anything. I just wouldn’t call LLM as exhibiting intelligence. It is great at making connections based on probability but doesn’t have a semantic understanding of what it is doing)
- stevenhuang 4mo agoYou do realize modern neuroscience considers the human brain as "just" a probability engine and that intelligence may well be the ability for an organism to predict well. > doesn’t have a semantic understanding of what it is doing I hope you realize this is an area of open, active research.
- Chu4eeno 4mo agoDidn't neuroscience some big scandals about bad statistics and overstating their findings (in addition to normal issues like replication)? Look up at least the "dead salmon study" (hint: it's related to fMRI, and you can probably guess its conclusions from its nickname). The "Voodoo Correlations" and "Cluster Failure" papers are also a bit eye-opening. In general we (humans) need to be humble about the limitations of our knowledge about how we function, it's an insanely complicated problem.
- jmalicki 4mo ago> In general we (humans) need to be humble about the limitations of our knowledge about how we function, it's an insanely complicated problem. We do. Which is why we shouldn't be assuming we're more than just probability engines, or be assuming we have more consciousness than a neural network.
- otabdeveloper4 4mo ago> to then make your model bigger, which then makes it smarter There's diminishing returns and at some point making a model bigger makes it dumber.
- fizx 4mo agoPerformance enhancements are what allow you to train a bigger model.
- gobdovan 4mo agoDeepSeek research: - V3 https://arxiv.org/abs/2412.19437 https://arxiv.org/abs/2412.19437 - V2 https://arxiv.org/abs/2405.04434 https://arxiv.org/abs/2405.04434 - R1 https://arxiv.org/abs/2501.12948 https://arxiv.org/abs/2501.12948 (RL applied to ML models was well-known beforehand, but they show it in the open, at scale, on big models) Then, there's the incentive analysis. If you can see that these models empirically get better with scale, why would you swap the main architecture? Those events will be pretty rare. I'm not saying there's noone cooking a new architecture, just that it is a pretty rare event. And it would have to come from some researchers that would be happy to not publish their findings, which is not really what a sizable portion of elite researchers (obviously not all) are incentivized to do. Of course, it's a bit of a verbal compression to claim simply 'scaled up'. They are recognisable scaled up transformers, but most new models come with a few tricks, but we're at the point where those usually are not an architectural rewrite and added to solve an explicit problem, like hallucination, not for big new capability gains.
- swyx 4mo ago> If you can see that these models empirically get better with scale, why would you swap the main architecture? Those events will be pretty rare c.f. hardware lotter https://arxiv.org/abs/2009.06489 https://arxiv.org/abs/2009.06489
- matusp 4mo agoThere are thousands of people working in top level labs. Somebody would leak it
- HarHarVeryFunny 4mo agoWe know for sure the architecture of the open weights models since llama.cpp understands the architecture it needs to build to plug the weights into to run them. It's always possible that the latest closed model is doing something architecturally different than the open weights ones we know about, but judging by how close the large open weight models such as DeepSeek are to SOTA performance, this seems unlikely. When OpenAI first came out with their near-mythical "Strawberry" (aka "o1") thinking model there was all sorts of speculation that they had made some sort of architectural breakthough, but then DeepSeek replicated the capability and published how they did it, proving that it was just better training, not any architectural change. There have been minor changes to the architecture over the years, but these are basically all efficiency tweaks such as various types of attention (some pioneered in the open by DeepSeek) that better scale to large context lengths, and the confusingly named "mixture of experts" architecture, but what's more notable really is how little the architecture has changed. The capability gains have been coming from better training and better data.
- gobdovan 4mo agoThe secret sauce is also having the necessary 'creativity' to not get ceased and desisted into oblivion and jail from all the copyrighted material you trained your model on. Btw, not making a moral judgement, [0] shows Michael and Dalton from YC discussing why Ilya Sutskever had to leave Google to pursue what's now ChatGPT [0] https://youtu.be/E8pvgN1j-Ck?t=748 https://youtu.be/E8pvgN1j-Ck?t=748
- root-parent 4mo agoThere is a whole moral judgement to be made here...lets hope Ilya wont get too pissed off if somebody leaks the work of his new initiative...information wants to be free and all that... Also would love to know if the same Legal team advised on Gemini...
- someguyiguess 4mo agoAnd to make anyone who threatened to expose them “commit suicide”
- miltonlost 4mo agoHe's a massive massive thief that people who have stolen far less from a convenience store have gone to prison for. The man is a villain.
- locknitpicker 4mo ago> The secret sauce though is all the datasets, RL training, knowledge of what works from doing all kinds of ablation experiments, and a massive compute moat. ReAct loops and tool-calling are the critical development feature. They turn a model from something that generates text into something that can independently influence the world around them. Without agent features, you have just a chatbot.
- galaxyLogic 4mo agoThe big breakthrough is we can interact with the agents using natural language - because of the LLM. It is the combination of LLM and agent-harnesses that make it look really smart. Agent-harness is a programmatic device that lets us tap into the vast knowledge in the LLM. It is probabaly true that many TV-commentators fail to appreciate this fact and therefore think LLMs are super-intelligent. No, it is the combination of LLM and the programmatic agent-haness that is the breakthrough. An interesting thought is that the LLM could in theory code the agent-harrness, start it running every time we interact with it. Currently the agent-harrness I think is pretty static I think. In theory it could be dynamically created for every task. Would that make it better don't know.
- locknitpicker 4mo ago> The big breakthrough is we can interact with the agents using natural language - because of the LLM. Without ReAct and tool calling, all you have is a chatbot. That's useful, but it's just a chatbot. ReAct loops and tool calling is what unblocks high value usecases. It enables systems to actually address free-form problem statements, gather data that is not a part of their training set, inspect the current state of services,and trigger actions in external systems. This goes well beyond mere chatbots. > It is the combination of LLM and agent-harnesses that make it look really smart. It's really not about "smart". It's about autonomous systems, and being able to consume and analyze new data, and trigger actions in external systems.
- Chu4eeno 4mo agoIt's not very novel, though, it's a fairly obvious step once you get something that can operate iteratively and largely independent, there were a ton of people trying to get LLMs to loop on their own even before deepseek r1. And I remember talking about goal directed behavior (which what people are calling "agents" now don't seem to properly have) and autonomous operation decades ago in the intelligent agent course at uni, including react loops. So no, the huge step with LLMs really was just that attention mechanism from that translation paper everyone forgot until Google brought its marketing to it, everything else is either just optimization/scaling, more money or old ideas suddenly relevant.