3 ms·if on MacOS I recommend llm-mlx which currently renders tokens 10%-15% faster than llama.cpp.by shironnnn_ 4mo agoif on MacOS I recommend llm-mlx which currently renders tokens 10%-15% faster than llama.cpp.