8 ms·
I’m getting better results with 30B unquantized (f16) than with 65B at 4 bits, FWIW. On a Mac Studio with 128GB RAM.
by enduser 4y ago
I’m getting better results with 30B unquantized (f16) than with 65B at 4 bits, FWIW. On a Mac Studio with 128GB RAM.
- wjessup 4y agopost example prompts and results please?
- koheripbal 4y agoThat's to be expected. 4 bits is too small. But the 8 bit 65b should outperform significantly. Have you tried it?
- sebzim4500 4y agoIs 4 bits too small? Or is the quantization method just not sophisticated enough. See https://arxiv.org/pdf/2210.17323.pdf https://arxiv.org/pdf/2210.17323.pdf
- enduser 4y agoI’ll give it a try this evening.
- sebzim4500 4y agoFrom what I can tell, the quantization being used there is extremely naive. If they were using techniques like GPTQ, presumably it would work much better. https://arxiv.org/pdf/2210.17323.pdf https://arxiv.org/pdf/2210.17323.pdf