7 ms·
More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/
by Tiberium 3mo ago
More details:
- https://platform.kimi.ai/docs/guide/kimi-k3-quickstart https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
- https://platform.kimi.ai/docs/pricing/chat-k3 https://platform.kimi.ai/docs/pricing/chat-k3
1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified.
This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5 which is currently on discount), and very close to 5.6 Terra pricing (Terra's input is $2.5).
One thing to consider, though: reasoning efficiency matters directly for how expensive a model actually is in real use. GPT's models are extremely reasoning efficient, and some Claude models like Fable at lower effort are as well. So if Sol spends 10K reasoning tokens to do something (at $30/1M) vs Kimi K3 that spends 50K reasoning tokens, Sol would win on cost effectiveness.
- Deukhoofd 3mo agoI feel like the quickstart is missing something. It's referring to its tech blog for actual benchmarks, but K3 isn't mentioned on there, the last thing on that blog was K2.6, 2 releases ago.
- gruez 3mo ago[dead]
- dghlsakjg 3mo agoTokenizers also matter. Anthropics tokenizers will encode the same piece of text at a way higher token count than OpenAi, for example. That said, Kimi is competing against GLM in my mind, and GLM 5.2 is less than 1/3 the price.
- asenna 3mo agoWith that kind of pricing, I don't think they're competing with GLM with this new launch.
- leecommamichael 3mo agoTokenizers define the alphabet on which the language model is trained. I don't want people to get the impression it's a module which can be swapped out or modified on its own. Alphabet size is a design consideration related to correctly encoding the training data.
- smallerize 3mo agoThat's true, but it makes it difficult to compare pricing when it's based on tokens. Maybe we need a benchmark for price per a specific input, like enwiki8.
- leecommamichael 3mo agoYes, almost all work people share which seeks to measure the capabilities and differences of models needs to get more precise. We are clamoring to say something meaningful about these things.
- victorbjorklund 3mo agoIt is kind of a shame we ended up comparing token pricing across models and providers when it doesn’t really make sense. Not sure what would be better though.
- whoopdeepoo 3mo agoWell isn't that what benchmarks are for? They compare total cost for a unit of work.
- alain94040 3mo agoUse price per page (standard English text)? That would also help make the metric easier to visualize. If you think a page is too vague, use a famous known writer's work as a reference.
- whodatbo1 3mo agoA better metric is price per byte. Most thinking traces, prompts, skills are in plain English, which is roughly 1 byte per character, assuming UTF-8 encoding (even code should not be much more either). As an aside, it is common to use bits-per-byte as a loss metric instead of the per token calculation, precisely because of the effect of different tokenizers.
- cmrdporcupine 3mo agoGLM is actually quite expensive in actual practice because it's not very token efficient. I've yet to find a way to run it on a monthly sub reliably for cheaper than Codex. Neuralwatt was cheap (but slow) but they cranked their price. Ollama monthly sub is speedy but doesn't offer a lot of quota. Right now unless you're paying by the token, there's no cost based reason to use the open weight models for daily coding work because the monthly coding plans from Anthropic and OpenAI are a better deal.
- computerex 3mo agoI know GLM is relatively expensive and so is Kimi, in comparison to those DeepSeek V4 pro and flash are a godsend and are absolutely good value.
- arizen 3mo agoAnd DeepSeek V4 Flash + GLM 5.2 is a really good blend of both (fast/cheap DS + more intelligent GLM)
- calgoo 3mo agoExactly, I design in glm 5.2, and then build with deepseek pro. I pay deepseek by token count, and honestly I added $20 and it just keeps going, hardly burning anything.
- versteegen 3mo agoI think MiMo 2.5 seems to be better than DS4 Flash, at the same price. DS4F can write pretty advanced code but it way overthinks the simple stuff, its CoT is full of errors (immediately corrects itself), and yesterday I was shocked it reliably struggled to spot a misplaced % in a format string when reading a whole file.
- pimeys 3mo agoI use V4 flash as my personal agent. It categorizes documents, organizes my calendar, searches information etc. for pennies. Amazing model. Not very good for programming though.
- mdasen 3mo agoIt also depends on how many tokens it needs to burn through to accomplish something. At this point, I always look at things like Artificial Analysis' total cost to run their tests. It'll take into consideration the cost of tokens, how many tokens it burns through, and how effectively it uses caching (and the price of that caching). If a model "costs the same" but its reasoning ends up going through a ton more tokens, it doesn't really cost the same in real world usage.
- KennyBlanken 3mo agoPrecisely. GLM 5.2 Thinking is pretty damn good - but it regularly does something nonsensical. Or even spits out what looks like a fragment of its memory cache. Or returns a bunch of Chinese. I find myself having to resubmit a query very often...so it being a third of the cost of other AIs isn't really relevant.
- zvikara 3mo agoI believe Kimi is spending more on marketing than GLM (a lot of ads lately) so I guess that's part of what the higher price supposed to cover.
- InsideOutSanta 3mo ago> That said, Kimi is competing against GLM in my mind, and GLM 5.2 is less than 1/3 the price. Having used GLM 5.2 extensively and K3 for a few hours now, these models are nowhere near each other. 5.2 is a great model, and I use it for a lot of things, but it's noticeably below Opus 4.8 or GPT-5.5 in real-world usage. K3 is in the same ballpark as Fable or Sol.
- wmedrano 3mo agoWhere does the less than a third the price come from? From the provided benchmark suite, it came slightly under half.
- deleted 3mo ago[deleted]
- martinald 3mo agoWill be interesting to see how it stacks up pricing wise on the various inference providers.
- sixtyj 3mo ago[flagged]
- shrubby 3mo agoSadly these days this seems like the least worse of the three major regimes.
- lostmsu 3mo agoYou are in a bubble. They just raided independent book stores in Hong Kong.
- em500 3mo agoEverybody is in a bubble. Which is why it's worth looking into other people's bubbles occasionally. https://www.pewresearch.org/global/2026/07/15/people-in-many-countries-now-view-china-more-positively-than-the-u-s/ https://www.pewresearch.org/global/2026/07/15/people-in-many...
- lostmsu 3mo agoWell no, that's why I build my opinion independently from other people.
- stavros 3mo agoI measure good and bad by proximity to me. China can directly hurt me the least, the US can hurt me the most.
- cyanydeez 3mo agoI eat 1M context in a local model in about 3-4 hours. It'd need to be exceptionally smart and error free to ever make sense.
- mmaunder 3mo agoAgreed re reasoning. I’ve seen this play out with 5x reasoning negating cost savings.
- schmorptron 3mo agoAre thinking models only the reasonable tradeoff vs using much larger non thinking ones because the cost of output tokens is below that of input tokens?
- csomar 3mo agoIt seems the subsidized era is nearing its end and we'll see a convergence on API pricing before a pulling of subscriptions pricing.
- easygenes 3mo agoThat’s not what this indicates. This is the biggest and most expensive to serve, and most capable open weights model yet. They’re just pricing it in line with capabilities. Kimi also offers generous subscriptions. Subs aren’t going anywhere. Think of subs like running an insurance business. There might be some users you lose money on (ones who max out their weekly quota without fail), but they’re managed such that the average subscription turns a healthy profit. There’s never been subsidies in model serving, inference is just cheaper in terms of ops TCO than people assume, and API margins are very high.
- csomar 3mo ago> They’re just pricing it in line with capabilities. So... convergence? > but they’re managed such that the average subscription turns a healthy profit. It didn't work like that, or at least that's not how it played out. People max-out their subs all the time which is why strict and multiple limits were implemented by all providers. Also, I subscribe to z.ai and recently they dropped the quota significantly that now their sub offers less than Claude and OpenAI. It's still x5-6 what it would cost on API costs though. > inference is just cheaper in terms of ops TCO than people assume, and API margins are very high. API margins (at least american ones) are probably healthy. But I don't think that inference is that cheap. It would cost 300-500k to just run GLM 5.2. There are lots of other factors too: reliability (can you keep the GPUs running all time), electricity cost, sys. admin costs, location costs, etc.. I wouldn't be surprised if the API margins are quite close to operational costs.
- ilaksh 3mo agoIt's as good as gpt 5.6 sol and _half_ the cost..
- nullbio 3mo agoAh, the old "subsidized" meme always rearing its head. Yawn.
- h14h 3mo ago> reasoning efficiency matters directly for how expensive a model actually is in real use I have high hopes on this topic, given token efficiency seemed to be the primary (only?) goal of the K2.7 Code release. Excited to see the signals that come out of the big eval/benchmark sites.
- sroerick 3mo agoHow do Kimi's subscriptions work? I find their price structure pretty confusing
- hedora 3mo agoDoes it have safety guardrails that constantly false positive like Claude does? The only obvious change I’ve seen since opus 4.6 came out is that it constantly flags my requests (no, I’m not doing biology research or security research, yes, it flags for both of those things). Recently, they backported the blocks to Opus 4.8, so I’m reluctantly stuck on sonnet. I probably could successfully apply to get special approval to use claude code unencumbered, but I don’t think it is ethical to support tooling that’s built so a central authority gets to decide what intellectual endeavors and knowledge work are permissible, and what are not.
- fmind-dev 3mo agoAPI prices are amazing, but hosting this on-premise will be real challenge.
- ImageXav 3mo agoI've been avidly using Fable since it was re-released and while it has been excellent at building the apps I want, the reasoning has been completely opaque. Kim, however, has exposed the whole reasoning trace, or enough of it to matter. I'd almost forgotten how nice it is to see this. I've been able to see all of the weird twist and turns it takes and it is joyful. But also, far, far more informative and means I can debug ideas far more thoroughly. Also, at a first glance it seems to have gotten quite far on a niche hobby horse of mine that no LLM has been able to crack. I'll be testing this more for sure.
- mahkeiro 3mo agoThe reasoning is key as most of the time the summary provided by fable is not enough to understand the choice and correct the logic. You have to either fully trust it or go to an exhaustive code review. This with the fact that you can only use 4.8 to security review the code produce by fable are the reasons I will not renew my anthropic subscription, the current experience is way to degraded.
- f3408fh 3mo agoWhat will you be replacing it with, if anything?
- epistasis 3mo agoI have severe complaints about Anthropic's product managers on this front. Their preference for hiding, obscuring, and trying to wrest control from the user are a bit harrowing. It would be wonderful to go back to Claude Code from before March. It seems like every release destroys value for me!
- qeternity 3mo agoIt's a defensive tactic to reduce the effectiveness of distillation. Say of that what you will, but it's not because they want to wrest control from users. It's because they don't want Chinese companies to do exactly what Moonshot (Kimi creators) and others have done.
- darkbatman 3mo agoalso its pretty big model inference costs are high even with margins running a 2.8T model costs a lot. if they release oss may be it goes down to $10-12 per million tokens.