2 ms·
I'd go for at least 32GB+. It'll fit in 24GB but leaves you little to no room for context, and that's at 4-bit quantization. If you want to run unquantized, yo
by thewebguyd 3mo ago
I'd go for at least 32GB+. It'll fit in 24GB but leaves you little to no room for context, and that's at 4-bit quantization.
If you want to run unquantized, you definitely need 128GB.
- Catloafdev 3mo agoNobody runs unquantized, there's literally no reason to. Q8 would be the largest anyone actually runs on consumer hardware for inference.
- deleted 3mo ago[deleted]
- bityard 3mo agoHalving the precision of the weights is not a free lunch...
- Catloafdev 3mo agoQ8 is virtually lossless. The quantization is much more noticeable around Q4 and below. FP16->Q8 on consumer hardware is 2x the speed at ~99.99% the quality.
- rvba 3mo agoAny source that confirms the 99.99% quality?
- Catloafdev 3mo agoI don't have a 'source' off-hand but I recommend reading up on it if you want to learn more. A lot of models on HF show a card demonstrating the different quality trade-offs between quants.
- gchamonlive 3mo ago[dead]
- bitexploder 3mo agoIt also comes down to inference speed, not "can I run this". 8-bit quant is quite a bit slower on an M5 Pro.