3 ms·
Cerebras is trialing Kimi K2.6 at 3000t/s (invite only). I'm excited for when the fast hardware gets more mainstream for frontier models. Models designed for sp
by scosman 4mo ago
Cerebras is trialing Kimi K2.6 at 3000t/s (invite only). I'm excited for when the fast hardware gets more mainstream for frontier models. Models designed for speed on Nvidia are nice addition that could bridge the gap.
- lostmsu 4mo agoCerebras currently does not provide any discounts for prefix caching making its use for agentic workloads sqr(n_turns) more expensive.
- michael-ax 4mo agonow that's what i call a software development breakthrough/platform! thanks for the heads up!
- adrian_b 4mo agoTFA mentions that until now special very expensive hardware like Cerebras was required for reaching this kind of speeds, and it emphasizes that what is novel in their results is that they have obtained over 1000 token/s for a model with over 1 T parameters by using just standard hardware, i.e. one server with 8 GPUs.
- btian 4mo agoSource? Their website says 1000t/s https://www.cerebras.ai/blog/which-is-faster-gemini-3-5-flash-or-kimi-k2-6-on-cerebras https://www.cerebras.ai/blog/which-is-faster-gemini-3-5-flas...
- scosman 4mo agoThis is likely correct, sorry for the bad info. Was working from memory.
- johndough 4mo agoCerebras got lucky that they IPOed last month instead of now.