4 ms·
> higher param count models will remain smarter for a looong time They're not smarter, they just know more stuff. You probably don't need knowledge about Poke
by otabdeveloper4 5mo ago
> higher param count models will remain smarter for a looong time
They're not smarter, they just know more stuff.
You probably don't need knowledge about Pokemon or the Diamond Sutra in your enterprise coding LLM.
The "smarts" comes from post-training, especially around tool use.
- anon7725 5mo agoIf the smarts came from post-training, we could show significant gains by doing that post-training again for previous generations of models. But we know that isn’t happening - effective post training is necessary but not sufficient for model performance.
- otabdeveloper4 5mo ago> we could show significant gains by doing that post-training again for previous generations of models That's what Chinese models are doing, and beating Opus et al.
- CamperBob2 5mo agoYou probably don't need knowledge about Pokemon or the Diamond Sutra in your enterprise coding LLM. That's one of the biggest remaining head-scratchers in this whole business. You do need all that unrelated stuff to make a good coding model. Nobody knows why you can't build a coding model by training on nothing but code, CS texts, specifications, and case studies, but so far it appears that you can't.
- otabdeveloper4 5mo agoThis one is kind of obvious - because people prompt coding LLMs with natural language. That's unrelated to stuffing the pre-train set with trivia factoids. An LLM that knows English very well isn't actually very large and certainly not hundreds of billions of parameters.