3 ms·
I don't find the authors points to be very convincing. There's a bit of a self-contradiction in the arguments being made -- namely that if you want better gener
by deepsquirrelnet 3y ago
I don't find the authors points to be very convincing. There's a bit of a self-contradiction in the arguments being made -- namely that if you want better generalization, you need more parameters (which equates to higher energy consumption in training).
However, the obvious retort here is that if you train a model that is good at generalizing, then you don't need train more models! A show of hands, who has used GPT or an open LLM vs who has trained one would yield a vast disparity. If you don't need generalization, you don't need huge models. Small models are efficient over narrow domains and don't require vast compute/energy resources.
Secondarily, it's a self-solving issue. Energy isn't cheap, and GPUs aren't cheap. If you're going to burn 10's of thousands of dollars in energy costs, you should probably have a decent reason to do it. But those reasons are quickly diminishing as things that have already been _done_.
Third, overparameterized models are becoming less of an issue during inference with efficient quantization techniques. Distillation, though harder, is another option. Again, you do can these things one time after training.
- ShamelessC 3y ago> A show of hands, who has used GPT or an open LLM vs who has trained one would yield a vast disparity. Believe me, that has almost nothing to do with "only needing one model" and everything to do with the compute required being absurd. > Small models are efficient over narrow domains and don't require vast compute/energy resources. The "best" small models are still typically born from the "best" overall models (however large) using a student/teacher paradigm. This is active research. As such, it is typically encouraged to go for "the big one", as you can distill it to far better small models than if you were to train the small models from scratch first. > If you're going to burn 10's of thousands of dollars in energy costs, you should probably have a decent reason to do it. But those reasons are quickly diminishing as things that have already been _done_. Again, this is _not_ the reason researchers don't train their own GPT-4's. They _absolutely_ would if they could raise that much money. The notion that researchers have everything they ever could have wanted from OpenAI's API-based walled garden approach is patently absurd. If they would release the weights, then finetuning and such could occur and you might have a point. To be clear, I'm not defending the article's main point. You just happened to have used some strange arguments against it (imo). The best argument is a sibling comment which points out that new architectures will be discovered that are more parameter efficient and run better on the hardware available. The burgeoning field of architecture search may mean we even have the ability to get the models themselves to help speed up this process of finding better ML architectures. It's always difficult to know when we have hit a wall of sorts however, and it may be that the transformer (and similar approaches) can only be refined so much. Time will tell.
- chaxor 3y agoOnce you have an NN model and you know you want to keep the model weight set, moving to an optical neural network can massively drop the energy use. Its not necessarily easy, and certain architectures may not be as amenable, but it can certainly be a path to reducing energy use.
- ShamelessC 3y agoI’m not aware of anyone having done this for a production ready scenario, I thought it was all still highly experimental and very early days?