3 ms·
> 5060 I think you're referring to a 5060Ti 16GB, yes? 32k context is easily done there. 64k can work with a more aggressive quant, but you lose a bit of spee
by fshr 26d ago
> 5060
I think you're referring to a 5060Ti 16GB, yes?
32k context is easily done there. 64k can work with a more aggressive quant, but you lose a bit of speed.
- kennywinker 26d agoYeah 16gb. For longer context, a 3bit quant is needed. Yes it’s tight on a 16gb card - can’t wait for the bubble to pop so hardware prices fall. But I don’t quite follow you - how does a more aggressive quant slow it down? Less bits per token means faster inference not slower.