3 ms·
I didn't realise the 5070 is slower than the 3090. Thanks. If you want a bit more context, try -ctv q8 -ctk q8 (from memory so look it up) to quant the kv cach
by idonotknowwhy 1y ago
I didn't realise the 5070 is slower than the 3090. Thanks.
If you want a bit more context, try -ctv q8 -ctk q8 (from memory so look it up) to quant the kv cache.
Also an imatrix gguf like iq4xs might be smaller with better quality
- parched99 1y agoI answered the question directly. IQ4_X_S is smaller, but slower and less accurate than Q4_0. The parent comment specifically asked about the QAT version. That's literally what this thread is about. The context-length mention was relevant to show how it's only barely usable.