3 ms·
> I'm expecting LLMs with hundreds of billions and eventually trillions of parameters will be able to run locally on my laptop and mobile phone, in the not-too-
by fpgaminer 4y ago
> I'm expecting LLMs with hundreds of billions and eventually trillions of parameters will be able to run locally on my laptop and mobile phone, in the not-too-distant future
Perhaps. There's been a lot of focus on training-compute optimal models in the industry. Rightfully so, as proofs of concept. That's what led to this perceived parameter count race in published models.
But remember the other side of the scaling laws. For inference, which is what we want to do on our phones, it's better to be inference-compute optimal. That means smaller models trained for longer.
As far as we know today there are no limits of the scaling laws. A 1B parameter model _can_ beat a 1T parameter model, if trained for long enough. Of course it's exponential, so you'd have to pour incalculable training resources into such an extreme example. But I find these extreme examples elucidating.
My pet theory these days is that we'll discover some way of "simulating" multiple parameters from one stored parameter. We know that training-compute optimal models are extremely over-parameterized. So it isn't the raw capacity of the model that's important. It seems like during training the degrees of freedom is what allows larger models to be more sample efficient. If we can find a cheap way of having one parameter simulate multiple degrees of freedom, it will likely give us the ability to gain the advantages of larger models during training, without the inference costs later.
I don't disagree that we're likely to see more and more parameter capacity from our devices. I'm just pointing out that the parameter count race is a bit of an illusion. OpenAI discovered the scaling laws and needed a proof of concept. If they could show AI reaching X threshold first, they could capture the market. The fastest way to do that is to be training-compute optimal. So they had to scale to 175B parameters or more. Now that it's proven, and that there's a market for such an AI, their and other's focus can be on inference-optimal models which are smaller but just as smart.
- cs702 4y ago> I don't disagree that we're likely to see more and more parameter capacity from our devices. I'm just pointing out that the parameter count race is a bit of an illusion. OpenAI discovered the scaling laws and needed a proof of concept. If they could show AI reaching X threshold first, they could capture the market. The fastest way to do that is to be training-compute optimal. So they had to scale to 175B parameters or more. Now that it's proven, and that there's a market for such an AI, their and other's focus can be on inference-optimal models which are smaller but just as smart. Good point. That could very well be what they're thinking about, in addition to potential improvements in training data and RLHF methods. Also, I agree it would be great if anyone figures out how to do something akin to "making a smaller model act as if it were gigantic during training" OR "pruning a gigantic model's 'dead paths' as it learns during training," to get the benefits of scale in training without its costs at inference.
- ilaksh 4y agoIgnorant comment but my limited "understanding" is that in most networks much of the knowledge is highly entangled and to do an efficient computation for a particular completion might involve something like converting the input sequence into a representation that involves the right "modular" latent sub-spaces, only, for the meat of the computation. So making those latent sub-spaces might be key. Although unfortunately I don't really understand any of it. But I strongly suspect that things need to be better factored if we want good interpretability and efficiency. Not that I think it's easy to do that and still get performance and generality.
- cs702 3y agoDeep learning is still a trade. No one knows or understands much. What little we do know, we know from empirical evidence, obtained at great cost in terms of blood, sweat, and tears. And also lots of data and compute. Lots and lots of data and lots and lots of expensive compute.