3 ms·
Because OP is running it on an M3 Ultra with Ollama. He'll get much better perf with llama.cpp or MLX, both of which Ollama wraps, albeit very poorly.
by woadwarrior01 1mo ago
Because OP is running it on an M3 Ultra with Ollama. He'll get much better perf with llama.cpp or MLX, both of which Ollama wraps, albeit very poorly.