9 ms·
> Scaling applies to multiple dimensions simultaneously over time. A frontier model today could be replicated a year later with a model half the size Models "e
by usrbinbash 11mo ago
> Scaling applies to multiple dimensions simultaneously over time. A frontier model today could be replicated a year later with a model half the size
Models "emergent capabilities" rely on encoding statistical information about text in their learnable params (weights and biases). Since we cannot compress information arbitrarily without loss, there is a lower bound on how few params we can have, before we lose information, and thus capabilities in the models.
So while it may be possible in some cases to get similar capabilities with a slightly smaller model, this development is limited and cannot go on for an arbitrary amount of time. It it were otherwise, we could eventually make a LLM on the level of GPT-3 happen in 1KB of space, and I think we can both agree that this isn't possible.
> giving the LLM a harness that allows tool use like what coding agents have
Given the awful performance of most coding "agents" on anything but the most trivial problems, I am not sure about that at all.