4 ms·
The 33B model with 4 bit quantisation would be better. Check this PR out, you can see the chart showing that even the best 13B quantisation would be a far cry
by pocketarc 3y ago
The 33B model with 4 bit quantisation would be better.
Check this PR out, you can see the chart showing that even the best 13B quantisation would be a far cry from the 30B with 2 bit quantisation: https://github.com/ggerganov/llama.cpp/pull/1684 https://github.com/ggerganov/llama.cpp/pull/1684
- sa-code 3y agoThank you!