5 ms·
[Author] Completely disagree. Any analysis shows that you see perplexity reduction at 4 bits. Have a look at llama.cpp's results here: https://github.com/ggerg
by waleedk 3y ago
[Author] Completely disagree. Any analysis shows that you see perplexity reduction at 4 bits. Have a look at llama.cpp's results here:
https://github.com/ggerganov/llama.cpp#quantization https://github.com/ggerganov/llama.cpp#quantization
4 bit has a perplexity score 0.13 or so higher.
- mmoskal 3y agoWell, if you have a fixed RAM size, you're better off with the largest model you can fit at 4 bits (13B 4b is way better than 7B 16b despite being twice smaller).
- MacsHeadroom 3y agoYou're just wrong. You're looking at the wrong numbers. The perplexity score of a model with twice the parameters in half the bits (4bit) is FAR LOWER (ie better). If you are limited to X RAM and have two 16bit models of size 4X and 2X then the 4X model in 4bit will always be far superior to the 2X model in 8bit, with far lower perplexity. Compare 13B's 4bit perplexity of 5.3607 to 7B's 8bit perplexity of 5.9069. That is over 0.54 lower perplexity for the same RAM amount by using 4bit! That is MASSIVE!
- jiggawatts 3y agoAnother factor is that larger models degrade less when quantized. You have to wonder if running a huge model, say, 300B parameters at 2-bit quantization might be "optimal" in that it would fit into a single A100 or H100 GPU and likely outperform an 80B parameter 8-bit model...
- Ambix 3y agoNot sure here. The LLaMA models - yes, all weights fit in the small range between -2.0 .. 2.0 And some other models have more crazy numbers with even more crazier outliers within them, like you might have a weight of 12.00 between long array of typical small numbers around 0.00 I've read story about attempt to quantize RWKV model into the 4/5 bits which failed short due to the presence of outlier weights. The author told somewhere that bigger models had worse perplexity because of this.
- oh_sigh 3y agoA stupid question but...what about a 16x model in 1bit?
- renonce 3y agoHas binary neural network been implemented for Transformers yet?
- SethTro 3y agoAnother factor in favor of 8 bits is that inference speed will generally be 2x faster than a larger model at 4 bits.
- Taek 3y agoThere's also research showing that the perplexity reduction is less at higher parameter counts. E.g. a 65b parameter model barely has any impact at all when reducing from 16bit to 4bit
- Ambix 3y agoNot actually, you might see here how bigger models have much worse perplexity with 4bit due to the weight outliers: https://github.com/saharNooby/rwkv.cpp/issues/12 https://github.com/saharNooby/rwkv.cpp/issues/12 For LLaMA models - yeah, different story.