4 ms·Or maybe even middle class plebeian 24gb rigs?by orangepanda 2y agoOr maybe even middle class plebeian 24gb rigs?griomnib 2y agoAt that point just run 8b.pulse7 2y agoOr wait for the IQ2_M quantization of 70b which you can run very fast on 24GB VRAM with context size of 4096...griomnib 2y agoAt some point there’s so much degradation with quantizing I think 8b is going to be better for many tasks.