5 ms·
>Oh does llama.cpp use MLX or whatever? No. It runs on MacOS but uses Metal instead of MLX.
by irusensei 6mo ago
>Oh does llama.cpp use MLX or whatever?
No. It runs on MacOS but uses Metal instead of MLX.
- zozbot234 6mo agoANE-powered inference (at least for prefill, which is a key bottleneck on pre-M5 platforms) is also in the works, per https://github.com/ggml-org/llama.cpp/issues/10453#issuecomment-4148905254 https://github.com/ggml-org/llama.cpp/issues/10453#issuecomm...
- OkGoDoIt 6mo agoIs that better or worse?
- irusensei 6mo agoDepends. MLX is faster because it has better integration with Apple hardware. On the other hand GGUF is a far more popular format so there will be more programs and model variety. So its kinda like having a very specific diet that you swear is better for you but you can only order food from a few restaurants.