4 ms·
my understanding is that the engine used (pytorch transformers library) is still faster than llama.cpp with 100% of layers running on the GPU.
by rain1 3y ago
my understanding is that the engine used (pytorch transformers library) is still faster than llama.cpp with 100% of layers running on the GPU.
- itake 3y agoI only have an m1
- rain1 3y agoI don't think the integrated GPU on that supports CUDA. So you will need to use CPU mode only.
- itake 3y agoYep, but isn’t there an integrated ML chip that makes it faster than cpu? Or does llama.cpp not use that?
- rain1 3y agounfortunately that chip is proprietary and undocumented, it's very difficult for open source programs to make use of. I think there is some reverse engineering work being done but it's not complete.
- qeternity 3y agoIt's the Huggingface transformers library which is implemented in pytorch. In terms of speed, yes running fp16 will indeed be faster with vanilla gpu setup. However most people are running 4bit quantized versions, and the GPU quantization landscape as been a mess (GPTQ-for-llama project). llama.cpp has taken a totally different approach, and it looks like they are currently able to match native GPU perf via cuBLAS with much less effort and brittleness.