2 ms·
I have to disagree. Nvidia spent a lot of effort on researching improved numerical representations. You can see a summary in this talk: https://www.youtube.com
by cpldcpu 2y ago
I have to disagree. Nvidia spent a lot of effort on researching improved numerical representations. You can see a summary in this talk:
https://www.youtube.com/watch?v=gofI47kfD28 https://www.youtube.com/watch?v=gofI47kfD28
A lot of their work was published but went by unnoticed. But in fact the majority of their performance increase in new architecture is resulting from this work.
Reading between the lines, it seems that they came to the conclusion that a 4 bit representation with a group exponent ("FP4") is the most efficient representation of weights for inference. Reducing the number of bits in weights has the biggest impact on LLMs inference, since they are mostly memory bound. At these low bit numbers, the impact of using multiplication or other approaches is not really significiant anymore.
(multiplying a 4 bit wight with a larger activation is effectively 4 additions, barely more than what the paper proposes)