4 ms·
(I assume this question is about whether models need all those layers during training, even if they don't need them during inference) Yes. There's the so calle
by ImprobableTruth 3y ago
(I assume this question is about whether models need all those layers during training, even if they don't need them during inference)
Yes. There's the so called "lottery ticket hypothesis". Essentially the idea is that large models start with many randomly initialized subnetworks ("lottery tickets"0 and that training finds which ones work best. Then it's only natural that during inference we can prune all the "losing tickets" away, even though we need them during training.
It's kind of an open question how large this effect is though. As the article mentions, if you can prune a lot away, this could also just mean that the network isn't optimally trained.
- zoogeny 3y agoI figured this must be a well-known property of neural networks. I'll do some reading on the lottery ticket hypothesis. That is almost exactly what I was thinking when reading the article: sure after you have trained it you can prune the layers that aren't used. But I wasn't sure you could know/guess which layers will be unused before you train. It strikes me as an interesting open question since if it is the case that you need big networks for training but can use significantly smaller "pruned" networks for inference there are many, many reasons why that might be true. Determining which of the possible reasons is the actual reason may be a key in understanding how LLMs work.