5 ms·
>Finding that 70% of attention heads and 20% of feed-forward networks can be excised with minimal effect on in-context learning suggests that large language mod
by reqo 3y ago
>Finding that 70% of attention heads and 20% of feed-forward networks can be excised with minimal effect on in-context learning suggests that large language models are undertrained.
I thought that it was common knowledge that LLMs are undertrained, none of the publicly available loss graphs show any sign of convergence!
- novaRom 3y agoIt's a disadvantage of current SOTA models: they are easy to train, but they must be large wasting lots of weights in order to generalize well. Maybe another architecture, transformer's successor will be more economical - having less weights with more skills and knowledge.
- two_in_one 3y agoI think I've seen somewhere years ago an article claiming size is needed for training. Which means probably that after training model can be optimized to minimize the size. Purging? However, smaller model cannot be trained that well. From my experience with image processing the bigger the better, and, they all have their limits. Nothing new here.
- k__ 3y agoCan we use unoptimized to train optimized ones?
- versteegen 3y agoMaybe that's why the Phi models do so well for their size. I'm guessing they may have been trained close to convergence, but the loss graphs aren't published. Phi-1.5 (1.3B parameters) was trained on 150B tokens (5 epochs), yet phi-1.5-web was trained on 300B so they didn't stop for lack of compute. Phi-2 (2.7B params) was trained on 1.4T tokens, epochs unknown.
- littlestymaar 3y agoI dont think that's the case, tinyllama has been trained on much more data already (2.5T for 1.1B params, and they aim for 3T) and while it didn't show signs of convergence, it also has much worse results than the Phi models. The main takeaway of Microsoft models seems to be the one they got in Textbooks are all you need, that is: data quality gets you very far, even if you use artificial data.
- Legend2440 3y agoEspecially the OPT models they study here were known to be undertrained. It would be interesting to compare to a more modern model like llama 2.
- jjallen 3y agoSo why not train them more? Is it to save costs?