5 ms·
Does quantizing the models reduce their "accuracy"?
by zapdrive 4y ago
Does quantizing the models reduce their "accuracy"?
- MacsHeadroom 4y agoYes, but only minimally. Not enough for any human to notice. However, even this minimal amount can be avoided with GPTQ quantization which maintains uncompressed fp16 performance even at 4bit quantization with 75% less (video)memory overhead. References: https://arxiv.org/abs/2210.17323 https://arxiv.org/abs/2210.17323 - GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers [Oct, 2022] https://arxiv.org/abs/2212.09720 https://arxiv.org/abs/2212.09720 - The case for 4-bit precision: k-bit Inference Scaling Laws [Dec, 2022]