2 ms·
>Raw parameter counts stopped increasing almost 5 years ago, and modern models rely on sophisticated architectures like mixture-of-experts, multi-head latent at
by kgeist 6mo ago
>Raw parameter counts stopped increasing almost 5 years ago, and modern models rely on sophisticated architectures like mixture-of-experts, multi-head latent attention, hybrid Mamba/Gated linear attention layers, sparse attention for long context lengths, etc.
Agree, I recently updated our office's little AI server to use Qwen 3.5 instead of Qwen 3 and the capability has considerably increased, even though the new model has fewer parameters (32b => 27b)
Yesterday I spent some time investigating it:
- Gated DeltaNet (invented in 2024 I think) in Qwen3.5 saves memory for the KV kache so we can afford larger quants
- larger quants => more accurate
- I updated the inference engine to have TurboQuant's KV rotations (2026) => 8-bit KV cache is more accurate
- smaller KV cache requirements => larger contexts
Before, Qwen3 on this humble infra could not properly function in OpenCode at all (wrong tool calls, generally dumb, small context), now Qwen 3.5 can solve 90% problems I throw at it.
All that thanks to algorithmic/architectural innovations while actually decreasing the parameter count.
- Vachyas 6mo agoWhat you described sounds plausible (expected, even). But >Raw parameter counts stopped increasing almost 5 years ago Really? 5 years ago? Until just about 3 years ago OpenAI's latest offering was only ChatGPT 3.5 Most of the models people talk about now didn't even exist 3 years ago let alone 5. Even now, I don't know if parameter count stopped mattering or just matters less For example, I have no idea if the new Mythos is MoE but I'm pretty sure it's more parameters.
- kgeist 6mo agoI agree the original poster exaggerated it. But generally models indeed have stopped growing at around 1-1.5 trillion parameters, at least for the last couple of years. >Even now, I don't know if parameter count stopped mattering or just matters less Models in the 20b-100b range are already very capable when it comes to basic knowledge, reasoning etc. Improving the architecture, having better training recipes helped decrease the required parameter count considerably (currently 8b models can easily beat the 175b strong GPT3 from 3 years ago in many domains). What increasing the parameter count currently gives you is better memorization, i.e. better world knowledge without having to consult external knowledge bases, say, using RAG. For example, Qwen3.5 can one-short compilable code, reason etc. but can't remember the exact API calls to to many libraires, while Sonnet 4.6 can. I think what we need is split models into 2 parts: "reasoner" and "knowledge base". I think a reasoner could be pretty static with infrequent updates, and it's the knowledge base part which needs continuous updates (and trillions of parameters). Maybe we could have a system where a reasoner could choose different knowledge bases on demand.