2 ms·
The gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF
by quantumleaper 2mo ago
The gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF ecosystem provides a better selection of quants, in my experience Unsloth ones are excellent.
- MrScruff 2mo agoI thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.
- quantumleaper 2mo agollama-server also supports saving the kv cache to SSD. I had no issues with cache invalidation using pi.
- gcr 2mo agoTIL! When was this functionality added? It wasn’t in llamacpp when I looked in June