4 ms·
Question for those using it. Can the 7B really be used locally on a card with only 16GB VRAM? LLM Explorer says[1] it requires 15.4GB. That seems like cutting i
by guerrilla 2y ago
Question for those using it. Can the 7B really be used locally on a card with only 16GB VRAM? LLM Explorer says[1] it requires 15.4GB. That seems like cutting it close.
1. https://llm.extractum.io/model/Qwen%2FQwen2.5-7B,58qKLCI6aniiqWi9q1tMUj https://llm.extractum.io/model/Qwen%2FQwen2.5-7B,58qKLCI6ani...
- johndough 2y agoI am happily using qwen2.5-coder-7b-instruct-q3_k_m.gguf with a context size of 32768 on an RTX 3060 Mobile with 6GB VRAM using llama.cpp [2]. With 16GB VRAM, you could use qwen2.5-7b-instruct-q8_0.gguf which is basically indistinguishable from the fp16 variant. [1] https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-GGUF https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-GGUF [2] https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp