5 ms·
I get 80 tok/s on the same model, and it's pretty usable. I'm not chatting with it; it's either given a bag of tokens to generate an answer or it's doing some a
by 0x457 1mo ago
I get 80 tok/s on the same model, and it's pretty usable. I'm not chatting with it; it's either given a bag of tokens to generate an answer or it's doing some agentic coding.
https://openrouter.ai/anthropic/claude-opus-5 https://openrouter.ai/anthropic/claude-opus-5 is it worthless because its 65 tps?
re: gemini
https://openrouter.ai/google/gemini-3.7-flash https://openrouter.ai/google/gemini-3.7-flash worthless as well?
That being said, 14 tok/s is pretty slow.
- ActorNightly 1mo agoThe standard is not what you can do with it, the standard is what is the free alternative. Local LLMs need to be able to beat the rate limits on all the free models like Gemini to be useful. The only way to do that is to have high enough tok/sec, especially for agentic loops.