31 ms·
It is very expensive to train these base models so a smaller size is more practical if you aren’t a big company with hundreds of powerful GPUs at hand. Table 15
by pythux 3y ago
It is very expensive to train these base models so a smaller size is more practical if you aren’t a big company with hundreds of powerful GPUs at hand. Table 15 from LLaMA paper[1] has some insightful figures: it took 135,168 GPU hours to train the 13B version and a bit more than 1M GPU hours for the 65B version. And we are talking about A100 80GB GPUs here (expensive and scarce). Not everyone can afford these kinds of trainings (especially if it takes a few attempts; e.g. if you’ve got a bug in the tokenizer)
[1] https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf
- andreygrehov 3y agoHold on, are you saying I can grab the 13B OpenLLaMa model and train it? I thought all of these models are already pre-trained and represent sort of the end state. Am I completely missing the point?
- JohnKemeny 3y agoA neural network is just a bunch of weights. You can always continue modifying the weights as you see fit. A network is never "done" learning.