3 ms·
A 4-bit quantized 33B parameter model will fit on your GPU and you'll be able to use a 2048 token context too. (4-bit quantized larger models are better than sm
by TheUninformed 3y ago
A 4-bit quantized 33B parameter model will fit on your GPU and you'll be able to use a 2048 token context too. (4-bit quantized larger models are better than smaller 8bit/16bit models)
You can run 4-bit quantized 65B models on your cpu, but it is slow, 1-2 tokens a second instead of 8-15 people typically get with a gpu, but you need two 24gb or an enterprise card with 48gb of ram to load them there.
https://old.reddit.com/r/LocalLLaMA/wiki/models https://old.reddit.com/r/LocalLLaMA/wiki/models has the information you are irritated about not being listed.