7 ms·
I can actually run the entire Q4_K_S version of this in the gpu with my 3060, it's blazing fast too in this mode (~10 tokens pr second) with the latest llama.cp
by tyfon 3y ago
I can actually run the entire Q4_K_S version of this in the gpu with my 3060, it's blazing fast too in this mode (~10 tokens pr second) with the latest llama.cpp, should be the same for kobildcpp too.
- knaik94 3y agoI am confident my bottleneck is thermals and hardware related. I had the same speed when comparing koboldcpp and llamacpp a few versions ago. I am running it on a laptop on a full windows 10 install. It's a 10th gen i7 hex, 16g ddr4, and 2070 maxq mobile. In that context, I consider it remarkably fast.