3 ms·
How much more work is it to get those up and running?
by ComputerGuru 3y ago
How much more work is it to get those up and running?
- idonotknowwhy 3y agoAlmost none of you already have python. Download exl2, exui from github and run a few terminal commands. This let's me run the 120b param models, which won't fit in vram if I use llamacpp
- Rastonbury 3y agoWait, so using large models isn't limited by VRAM anymore?
- idonotknowwhy 3y agoIt is. I have 48GB of VRAM. But exl2 is more efficient, and can be quantized to partial bits. So you can run things like 4.75bpw, etc. I can run 120b models at 3bpw. The larger models like this are less affected (increased perplexity) by the quantization.
- ComputerGuru 3y agoDid you have to quantize it yourself to 4.75bpw and 3bpw or are they readily available for download?
- idonotknowwhy 3y agoMost of the time it's readily available eg: Panchovix/goliath-120b-exl2 (there's a different branch for each size) Some of them I've had to do myself eg. I wanted a Q2 GGUF of Falcon 180b There's a guy on huggingface called "TheBloke" who does GGUF, AWQ and GPTQ for most models. For exl2, you can usually just search for exl2 and find them.
- ComputerGuru 3y agoThanks, friend!
- idonotknowwhy 3y agoTypo: I meant "almost none if you already have python installed"