3 ms·
I run it locally at q4_k_xl on a r9700 with kv cache bf16 and while it thinks a lot, it’s still fast enough to do the task. This model had its knowledge replac
by pyrolistical 1mo ago
I run it locally at q4_k_xl on a r9700 with kv cache bf16 and while it thinks a lot, it’s still fast enough to do the task.
This model had its knowledge replaced with reasoning ability. The chain of thought what makes this reasoning effective.
So this is why you need to let it think and don’t quantize the kv cache.
- FeepingCreature 29d agoOr at least use a modern llama.cpp with KV activation rotation.