3 ms·
It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M
by MeProtozoan 2mo ago
It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M Output $15.00/M)
- tidbeck 2mo agoComparing to Opus 5: Claude Opus 5 (Uncached Input $5/M Cached Input $0.50/M Output $25/M) but you also pay a premium on Cache write 25% for 5m and 100% for 1h.
- chrsw 2mo agoThen there’s the questions of token efficiency and token quality.
- fnordpiglet 2mo agoI have to say cc opus 5 is abysmal. It talks to itself incessantly, gets stuck in minutia, fails to understand problems clearly and makes steering mistakes constantly. It also has a weird behavior where it says “ok I know exactly what to do and I will start now,” then sits waiting for user input. If you’re not on the ball you’re constantly losing 5m/1h cache. Just give me back 4.6.
- veber-alex 2mo agoI still see Opus 4.6 available in CC. Also I have been using Opus 5 for the last few days and I find it work fine. It's a little verbose but the code quality is good.
- dannyw 2mo agoYou can even use /model to use Opus 4.5 if you want. Broadly though I find Opus 5 and most of their point releases (4.7 being the exception) to be excellent and upgrades.
- serial_dev 2mo agoFable ends up in a confusion loop for me all the time.
- btown 2mo agoFireworks' priority tier of Kimi (at $3.75/M vs. Moonshot's $3.00/M) is available on OpenRouter as well. https://openrouter.ai/moonshotai/kimi-k3#providers https://openrouter.ai/moonshotai/kimi-k3#providers Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!
- pimeys 2mo agoI've used GLM-5.2 a lot on fireworks and had never ever issues on rate limits. If they cannot handle the load with K3, there's the priority tier to get your evals done. I'm definitely having full eval suite on already if they get overloaded later on.
- btown 2mo agoSelf-replying as this is no longer correct - there are two tiers, one matching Moonshot pricing with comparable latency/throughput, and a fast mode at $4.50/M (still cheaper than Opus) with 3x the throughput.
- lmf4lol 2mo agoI love fireworks.ai! They launched it couple of hours ago and we have it now already live on our platform for our users. Just a shame they deprecated the on-demand flux models :( Where do I get my fix for image gen now?
- dkersten 2mo agoYep, it’s also availability on Together.ai for your sale price. Fireworks and Together were the first places I checked!