4 ms·
What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases over
by kamranjon 2mo ago
What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
- beacon294 2mo agoTry the llama.cpp fork by thetom. It's called turboquant after the technique