3 ms·
Yeah, over 65GB VRAM... that'd be expensive but not impossible. I think three RTX 4090's could do it, with their 24GB each.
by guerrilla 2y ago
Yeah, over 65GB VRAM... that'd be expensive but not impossible. I think three RTX 4090's could do it, with their 24GB each.
- csomar 2y agoOnly 21.3GB is required.
- guerrilla 2y agoIs this wrong? It says 65.8GB. If it's wrong, what source should I be using instead? https://llm.extractum.io/model/Qwen%2FQwen2.5-Coder-32B-Instruct,6nvrT0uDyEPCowu5gDhQAA https://llm.extractum.io/model/Qwen%2FQwen2.5-Coder-32B-Inst...
- exe34 2y agothe ollama one is probably quantised.
- mistercheph 2y agoOllama's default quantization is q4_0 which is quite bad, but you can go to ollama's model page to see all the quantizations they have available, you can do e.g. "ollama run qwen2.5-coder:32b-instruct-q8_0" which will need ~35G + space for context
- tmikaeld 2y ago... with 4-bit quant and 4000 token context.
- mistercheph 2y agoYou really don't need models to fit into GPU's anymore for inference, e.g. llama.cpp is pretty good at partial GPU offload and I've gotten pretty fast results with only being able to fit ~30% of a model in VRAM and the rest in DDR5.