3 ms·
does your comment depend on the OS? I thought MLX has better performance on MacOS than llama.cpp
by itake 2mo ago
does your comment depend on the OS? I thought MLX has better performance on MacOS than llama.cpp
- quantumleaper 2mo agoThe gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF ecosystem provides a better selection of quants, in my experience Unsloth ones are excellent.
- MrScruff 2mo agoI thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.
- quantumleaper 2mo agollama-server also supports saving the kv cache to SSD. I had no issues with cache invalidation using pi.
- gcr 2mo agoTIL! When was this functionality added? It wasn’t in llamacpp when I looked in June