3 ms·
Hmm, perhaps I should switch to an MLX version… problem is, it took quite a bit of work to get llama-server (with llama.cpp) to serve my model(s) and allow requ
by jwr 2mo ago
Hmm, perhaps I should switch to an MLX version… problem is, it took quite a bit of work to get llama-server (with llama.cpp) to serve my model(s) and allow requests in non-thinking (default) and thinking modes.
But 30-40 tokens/s would make a big difference.