5 ms·
It’s worth noting this is the 7B model (nonquantized). You can get this running on pretty much any GPU with 8GB VRAM and above. You can run the 13B model but th
by syntaxing 4y ago
It’s worth noting this is the 7B model (nonquantized). You can get this running on pretty much any GPU with 8GB VRAM and above. You can run the 13B model but that would take two GPU or reducing FP16 to FP8 (I haven’t tried it myself). A single connection for chatgpt is rumored to require 8X A100.
- nico 4y agoIt makes me wonder if this trend will kill NVIDIA. At this pace we might not even need GPUs anymore.
- zamalek 4y agoQuantizing it to 8-bit basically eliminates its ability to write code.
- Taek 4y ago? All the research I've seen says quantization has basically negligible performance impact. My experience working with 65B 4bit has been great
- syntaxing 4y agoI agree with OP. Only the nonquantized models has given me good results too. I only have used the 7B and 13B. I don't have enough computing power to run 65B.