3 ms·
Can I run this locally with a 4090? If so, what’s the easiest way?
by PcChip 2y ago
Can I run this locally with a 4090?
If so, what’s the easiest way?
- vunderba 2y agoEven a 4-bit quant version of the MoE 8x22 is going eat ~80GB of VRAM. No way you're running this on a 4090 without setting it up as a GGUF to split between VRAM and regular RAM, and then you're going to have to deal with the low token rate as a result. Running WizardLM-2 70B or lower WizardLM-2 7B is much more feasible however. Look into Ollama: https://ollama.com/library/wizardlm2/tags https://ollama.com/library/wizardlm2/tags
- 0x008 2y agoThey are based on llama2 architecture which performs a lot worse.