4 ms·
what quant are you running for that rig? i've been running q4, not sure if I can bump that up to q5 across the board (or if it's worth it in general)
by BaculumMeumEst 2y ago
what quant are you running for that rig? i've been running q4, not sure if I can bump that up to q5 across the board (or if it's worth it in general)
- a_wild_dandan 2y agoI run q5s usually, since it's a 40% haircut on model size, with nearly no PPL loss. (Presuming an 8b native model like Qwen.)