4 ms·
I'd guess the slowness is mostly due to there currently being only one provider, Moonshot AI. And they are overwhelmed with demand. Let's judge the speed of th
by Slartie 2mo ago
I'd guess the slowness is mostly due to there currently being only one provider, Moonshot AI. And they are overwhelmed with demand.
Let's judge the speed of the model when its weights are released and every inference provider on the planet offers it, so demand can spread out a bit.
It's the same topic with token budget comparisons and subscription pricing - don't people understand that this doesn't really matter for open weights models? The pricing is going to be determined by the inference providers, and until they had a chance to evaluate the model on their infra and set token prices accordingly, one doesn't really have anything tangible to compare with other open models nor with closed ones.
- epolanski 2mo agoEven if hardware capacity increases, it seems clear it uses way more tokens, so I don't expect parity with other competitors on that front. On the other hand I expect K3 future refinements to be massive and more efficient.
- Slartie 2mo agoThe speed at which tokens are crunched, even on the same hardware, differs between models as well. Using more tokens is only a problem if they are processed at the same speed as with a comparison model.
- EwanToo 2mo agoUsing more tokens is a significant problem if you pay per token?
- Slartie 2mo agoNot if the price per token is significantly lower. Also this arm of the discussion was about speed, not price.
- locknitpicker 2mo ago> Using more tokens is a significant problem if you pay per token? Yoh have posts in this thread suggesting that Fable is 5x more expensive than Kimi K3.