3 ms·
Wait, can llama.cpp accelerate inference without full VRAM for the model? Or do you mean Llama 65B quantized to fit?
by idkyall 3y ago
Wait, can llama.cpp accelerate inference without full VRAM for the model? Or do you mean Llama 65B quantized to fit?
- brucethemoose2 3y agoYeah, it can split weights. Whatever fraction of the weights that don't fit into vram will be computed on the CPU (with reasonable speed). Additionally, prompt processing will work with large models even with low vram GPUs.