3 ms·
Not yet. Without pure Apple game or decent GPUs, even with a lot of RAM and threads, all you get is about 30-50 tokens/second, and that's thinking turned off. W
by moezd 4mo ago
Not yet. Without pure Apple game or decent GPUs, even with a lot of RAM and threads, all you get is about 30-50 tokens/second, and that's thinking turned off. Without these optimizations your model will have a field day with your MCPs, skills and agent descriptions and you will watch the paint dry before seeing the first output token. Local model serving means you have to fight for every token in your context window, which is quite opposite of what Claude/GPT/Copilot are pushing the industry towards.
- amarshall 4mo agoThinking doesn’t change output speed. Anthropic’s models are ~ 40–60 t/s median output speed.
- moezd 4mo agoDo you have access to Anthropic model weights to run them locally?
- amarshall 4mo agoNo, and having that is not required to know output speed nor the effect of thinking, so I don’t see the point in such a superfluous, indirect question. As for the question you’re likely asking: benchmarks that include speed across many models and providers available at various places e.g. https://artificialanalysis.ai/leaderboards/models https://artificialanalysis.ai/leaderboards/models