6 ms·
There is no question that quantization degrades quality. The GGUF R1 uses Q4_K_M, which, on Llama-3-8B, increases the perplexity by 0.18[0]. Many plots show inc
by espadrine 2y ago
There is no question that quantization degrades quality. The GGUF R1 uses Q4_K_M, which, on Llama-3-8B, increases the perplexity by 0.18[0]. Many plots show increasing degradation as you quantize more[1].
That said, it is possible to train a model in a quantization-aware way[2][3], which improves the quality a bit, although not higher than the raw model.
Also, a loss in quality may not be perceptible in a specific use-case. Famously LMArena.ai tested Llama 3.1 405B with bf16 and fp8, and the latter was only 2 Elo points below, well within measurement error.
[0]: https://github.com/ggml-org/llama.cpp/blob/master/examples/quantize/quantize.cpp#L46 https://github.com/ggml-org/llama.cpp/blob/master/examples/q...
[1]: https://github.com/ggml-org/llama.cpp/discussions/5063#discussioncomment-10906969 https://github.com/ggml-org/llama.cpp/discussions/5063#discu...
[2]: https://pytorch.org/blog/quantization-aware-training/ https://pytorch.org/blog/quantization-aware-training/
[3]: https://mistral.ai/news/ministraux https://mistral.ai/news/ministraux