4 ms·
You don’t really need to fit fully in memory. Memory requirement to train is ~6DP * precision Where D is number of tokens*mini batch size and P is number of p
by ivalm 3y ago
You don’t really need to fit fully in memory. Memory requirement to train is
~6DP * precision
Where D is number of tokens*mini batch size and P is number of parameters.
So if you want to fit fully into memory with a mini batch of 1, context window 32k, and 16 bit precision, that’s
144e12/6/32e3/2 = 375M param.
If you apply one token at a time then
144e12/6/2 = 12 T param
Ofc, in reality you have model parallelism as well…