7 ms·
4-bit quantization tends to come at the cost of output quality losses. https://github.com/ggerganov/llama.cpp/issues/9 https://github.com/ggerganov/llama.cpp/is
by helloericsf 2y ago
4-bit quantization tends to come at the cost of output quality losses.
https://github.com/ggerganov/llama.cpp/issues/9 https://github.com/ggerganov/llama.cpp/issues/9
- ssheng 2y agoQuality loss with quantization is expected. It seems like with GPTQ the loss is within acceptable range based on the perplexity score shown.