3 ms·
The model I have is q4_0 I think that's 4 bit quantized I'm running in Windows using koboldcpp, maybe it's faster in Linux?
by smiley1437 3y ago
The model I have is q4_0 I think that's 4 bit quantized
I'm running in Windows using koboldcpp, maybe it's faster in Linux?
- brucethemoose2 3y agoI am running linux with cublast offload, and I am using the new 3 bit quant that was just pulled in a day or two ago.
- smiley1437 3y agoThanks! I'll have to try the 3bit to see if that helps
- LoganDark 3y agocuBLAS or CLBlast? There is no such thing as cublast
- LoganDark 3y ago> The model I have is q4_0 I think that's 4 bit quantized That's correct, yeah. Q4_0 should be the smallest and fastest quantized model. > I'm running in Windows using koboldcpp, maybe it's faster in Linux? Possibly. You could try using WSL to test—I think both WSL1 and WSL2 are faster than Windows (but WSL1 should be faster than WSL2).
- smiley1437 3y agoI didn't know what WSL was, but now I do, thanks for the tip!