3 ms·
Nobody runs unquantized, there's literally no reason to. Q8 would be the largest anyone actually runs on consumer hardware for inference.
by Catloafdev 3mo ago
Nobody runs unquantized, there's literally no reason to. Q8 would be the largest anyone actually runs on consumer hardware for inference.
- deleted 3mo ago[deleted]
- bityard 3mo agoHalving the precision of the weights is not a free lunch...
- Catloafdev 3mo agoQ8 is virtually lossless. The quantization is much more noticeable around Q4 and below. FP16->Q8 on consumer hardware is 2x the speed at ~99.99% the quality.
- rvba 3mo agoAny source that confirms the 99.99% quality?
- Catloafdev 3mo agoI don't have a 'source' off-hand but I recommend reading up on it if you want to learn more. A lot of models on HF show a card demonstrating the different quality trade-offs between quants.