3 ms·
Yeah I don't think any of the labs have some secret sauce for intelligence either. It seems most of the advancements are still coming from hardware, making LLMs
by password54321 3mo ago
Yeah I don't think any of the labs have some secret sauce for intelligence either. It seems most of the advancements are still coming from hardware, making LLMs more efficient and throwing more compute and data at problems. And even those problems still require a lot of prompt engineering: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_prompt.pdf https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98...
- andy99 3mo agoThe secret sauce is training data. They’re not just taking advantage of more compute (which obviously is necessary but as mentions basically a commodity). They are paying billions to data labelers and making judgements about the nature of the training data they best need to make the product they want. This seems to get pushed aside as a minor point but it’s the primary differentiator of the big labs.
- password54321 3mo agoAs a I said, compute and data. But LLMs can be distilled, so even their data is not much of a secret sauce.
- reinitctxoffset 3mo agoI'm pretty sure at this point that Anthropic is training mixture models (at least in the heavy pre-train) and deploying them dense with explicit loss on thinking trace coherence. Having a thinking trace that is legible, coherent, and immediately implies the explicit turn output and/or tool use seems difficult if not impossible to reliably get from mixture models. I predict MoE is a transitional technology, it's got too many problems and the benefits are...kinda grandfathered into the dogma at this point.
- charcircuit 3mo ago>I predict MoE is a transitional technology While scaling laws hold (more weights = better), and time / financial costs are not trivial the incentives are in place to have MoE. MoE means you can have more weights without increasing the critical path of evaluating it. I am curious what you believe the problems with it that would cause people to prefer using less weights. I'm not following what you mean by MoE can't have legible thinkings trace or tool use when existing models with MoE can.
- reinitctxoffset 3mo agoWeights are not created equal: while interpretability is a young field the prevailing view at the moment is that MLP (hence experts) in a mixture model are substantially where dense encoding of factual information resides, attention is even less easily interpreted but it should be uncontroversial that temporal/sequential modeling occurs here. So it's more consistent with available empirics to say that an architecture can be characterized along a spectrum from fully dense to mixture (a sub spectrum) to Engram-style lookup, and the amount of model power allocated at this point or that will recover different performance profiles. By far the most stark example of how much performance in reasoning is left on the table is Qwen3.6-27B, which depending on the task, comparison model, and whose benchmarks you believe outperforms mixture models 15-60x larger in total parameter count. It's badly under-studied (in public) because of the paucity of modern dense models at the near frontier, but even that one data point pretty much rules out the cocktail party version of the Chinchilla-adjacent scaling thesis (which wasn't about modern MoE to begin with). The "Mixture of Parrots" work is a good jumping off point if you want to get a modern literature review.
- charcircuit 3mo ago>the prevailing view at the moment is that MLP (hence experts) in a mixture model are substantially where dense encoding of factual information resides Yes, because that's where all the parameters are. For reference in GLM 5.2 98% of the weights are for the experts. >The "Mixture of Parrots" work is a good jumping off point The paper shows increased performance on knowledge dependent task while having similar reasoning capabilities. This backs up what I was saying about how the weights unlock extra performance without increasing inference costs as much as a dense model would. >model power allocated at this point or that will recover different performance profiles While increasing the number of weights makes the model better, where those weights are does matter in how much better the model gets and also matter in regards to the cost of training / inference. Model design is a big set of trade offs and I see MoE as a useful tool that will survive in the trade off space. >reasoning is left on the table Even so, if there was 2 models with an equivalent amount of reasoning ability and priced the same would you rather pay for the one with narrow knowledge or wider knowledge. >because of the paucity of modern dense models at the near frontier, You don't need to be at the frontier to benefit from MoE. Even open source models that are behind the frontier, benefit from being able to host experts on different machines, and scale individual, commonly used experts separately from each other. On the other end with small models you are probably resource constrained so you want to maximize the tokens generated per second. This makes going for purely dense models niche like you are saying.
- int_19h 3mo agoMoE is just activating fewer weights per token than the whole model. It will continue to make sense for as long as compute is more expensive than memory (at scale).
- dominotw 3mo agoeven meta that sucks at doing anything is releasing frontier models. making an top ai is easier than making twitter clone( threads) if you have enough money.
- password54321 3mo agoI mean the problem with Threads was lack of user engagement. The same could possibly still be said about their models.
- dominotw 3mo agoyes ofcourse. But engagement needs strategy and execution.