3 ms·
That's not really a concern. If you have a trillion parameter 8-bit fp network or a trillion parameter 1.5-bit ternary network, based on the scaling in Microso
by kromem 2y ago
That's not really a concern.
If you have a trillion parameter 8-bit fp network or a trillion parameter 1.5-bit ternary network, based on the scaling in Microsoft's paper the latter will actually perform better.
A lot of the current thinking is that the nodes themselves act as superpositions for a virtualized network in a multidimensional vector space, so precision is fairly arbitrary for the base nodes and it may be that constraining the individual node values actually allows for a less fuzzy virtualized network by the end of the training.
You could still have a very precise 'calculator' feature in the virtual space no matter the underlying parameter precision, and because each parameter is being informed by overlapping virtual features, may even have less unexpected errors and issues with lower precision nodes.
- xwolfi 2y agoYup, your response makes me think they should just use a calculator, like everyone.
- imtringued 2y agoI don't know what you mean. They already use the GPU as a calculator.
- ben_w 2y agoI believe you four are talking about different things; the models are executed on very good "calculators" (if you want to call the GPUs that), but themselves are not very good at being used as calculators. LLMs are sufficiently good hammers that people see everything as a nail, then talk about how bad they are at driving screws.