4 ms·
>I'm running gpt2-xl (1.5B params) locally with KV caching at 120ms/token (vs. 450ms without caching). That seems very slow compared to llama cpp?
by cypress66 3y ago
>I'm running gpt2-xl (1.5B params) locally with KV caching at 120ms/token (vs. 450ms without caching).
That seems very slow compared to llama cpp?
- smpanaro 3y agoYeah, I believe it is. You trade off speed for lower power usage and CPU. 8 tokens/sec is usable though.