4 ms·
You typically want as much VRAM as possible for this type of application. For example, Llama has versions that take 32GB of VRAM, even after quantization (comp
by ClassyJacket 3y ago
You typically want as much VRAM as possible for this type of application.
For example, Llama has versions that take 32GB of VRAM, even after quantization (compression):
https://old.reddit.com/r/LocalLLaMA/comments/1806ksz/information_on_vram_usage_of_llm_model/ka72kgc/ https://old.reddit.com/r/LocalLLaMA/comments/1806ksz/informa...
There are smaller versions too however if you're VRAM constrained.