3 ms·
Georgi just added support for all models (13B/33B/65B) [0] LLaMA 65B can do ~2 tokens per second on my M1 Max / 64 gb ram [1] [0] https://twitter.com/ggergano
by lawrencechen 4y ago
Georgi just added support for all models (13B/33B/65B) [0]
LLaMA 65B can do ~2 tokens per second on my M1 Max / 64 gb ram [1]
[0] https://twitter.com/ggerganov/status/1634488664150487041 https://twitter.com/ggerganov/status/1634488664150487041
[1] https://twitter.com/lawrencecchen/status/1634507648824676353 https://twitter.com/lawrencecchen/status/1634507648824676353
- KVFinn 4y agoVery cool. I've seen some people running 4-bit 65B on dual 3090s, but didn't notice a benchmark yet to compare. It looks like this is regular 4-bit and not GPTQ 4-bit? It's possible there's quality loss but we'll have to test. >4-bit quantization tends to come at a cost of substantial output quality losses. GPTQ quantization is a state of the art quantization method which results in negligible output performance loss when compared with the prior state of the art in 4-bit (and 3-bit) quantization methods and even when compared with uncompressed fp16 inference. https://github.com/ggerganov/llama.cpp/issues/9 https://github.com/ggerganov/llama.cpp/issues/9
- magoghm 4y agoOn my M1 Ultra LlaMA 65B generates ~3 tokens per second (using 16 threads).