8 ms·
Kimi K2.7-Code: open-source coding model with better token efficiency
- yanis_t 4mo agoI was wondering how does Anthropic and likes keep competitive when Opus is ($5 / $25) 5x times more expensive compared to Kimi K2.6 ($0.7 / $3.4) or other Chinese models, while being only marginally better. My theory is that US enterprise just can't send data to Chinese and that's understandable, but is that "the moat"?
- yababa_y 4mo agoI want Opus to be only marginally better, but I do mostly research engineering and its ability to not fuck up my projects is absent. Every time my credits lapse I let kimi and composer2.5 have some play and it’s basically just an excuse for me to keep playing computer because when the oai/ant credits refresh I always need to spend hours recovering from the other models either misconceptions or boneheaded eng practices. Even when I only let it touch my web games…
- greenavocado 4mo agoYou have to revert to Opus 4.5 and 4.6. I bet you'll see a massive improvement based on what you're describing
- benjiro3000 4mo ago[dead]
- re-thc 4mo ago> My theory is that US enterprise just can't send data to Chinese Lots of US providers are hosting these “open source” models so doubt that’s the problem.
- DCKing 4mo agoThe moat right now is model performance and what that means for how many tokens and additional time you spend. I say this as a relatively frequent user of Kimi models and generally a big fan. But on not-yet-gamed benchmarks like DeepSWE, Kimi K2.6 is beaten soundly by Claude Sonnet 4.6 ($3 / $15) and even slightly by GPT 5.4 Mini ($0.75 / $4.50). There's no question Kimi models are very good for a lot of code tasks. They're the best quality open weight model. But to get similar overall outcomes as on Sonnet/Opus, on average you'll spend many more tokens and will have to do more managing of the model. You shouldn't look at price per token, you should look at how much you pay for the entire process.
- papersail 4mo agoI'm not sure I would put too much weight on DeepSWE as a benchmark, given that GPT-5.4-mini ended up close to Opus 4.6 there.
- DCKing 4mo agoAny benchmark is iffy and has weird results, but this is the best we got at the moment. Most people working with Opus and Kimi would likely tell you they're much further apart than the numbers that were quoted for Kimi K2.6, and DeepSWE seems to capture that gap better. One major thing DeepSWE has going for it is that all other benchmarks (including those quoted by MoonshotAI on this page) don't: the other benchmarks that are completely gamed. The benchmark answers are public and part of each model's training data. This benchmark may still be iffy, but at least it's not gamed.
- WarmWash 4mo agoSomehow the internet has also forgot that cheating to get ahead in China is basically a norm and expected behavior.
- DCKing 4mo agoAmerican labs also use gamed and cherry-picked benchmarks extensively. Anthropic used them in their Fable announcement and avoided DeepSWE because it doesn't beat GPT-5.5 in that one. Google's numbers for Gemini 3.5 Flash recently did not at all line up with people's subjective experience using these models, and this also happened with Gemini 3.1 Pro before it. Everybody has incentives to manipulate benchmark results to show their models in the best light.
- nullbio 4mo agoI think none of them having a defacto and high quality English focused cli is a big part of it. None of the Chinese models I've tried have worked well in opensource cli's. Granted, I've only tried a few, but still...
- freigeist79 4mo agoi use github copilot cli + openrouter + qwen 3.7 max and it's really much better than i expected (used to opus 4.7 at work)
- Bnjoroge 4mo agohuh? They all work great in omp/opencode unless you mean their own native clis like kimi code
- saratogacx 4mo agoI've been using charm's Crush with GLM for several months and it's been working great. I've only seen it shift to non-english once and it was already in a wonky state when it flipped.
- khuey 4mo agoI think most people who've tried them both would tell you Anthropic's models are more than marginally better than Kimi. Kimi and the other open source models may score well on SWE-bench or whatever but the gap is noticeable IMHO once you actually try to use them.
- Bnjoroge 4mo agoIt depends on what your task is and how precise your prompts are. Planning with fable or 4.8 and laying out the plan in step by step process and coding with mimo v2.5 pro or dsv4pro or qwen 3.7 max and doing a final review with 5.5 has worked really well for me for infra stuff.
- mnicky 4mo agoCoding with sufficiently precise plan takes almost all real work from the implementator, doesn't it? So it's not a fair comparison...
- efromvt 4mo agoI think the perception is that it is not 'only marginally better'; whether or not you specifically agree that perceived quality gap lets them differentiate on price. I'd further say that there are probably enough rational actors running evals out there that the marginally better is not pure vibes for the cases where people are spending lots of money, but I only have direct line of sight to some of those eval suites. Maybe everyone is irrational and anthropic is exploiting that!
- smoe 4mo agoI reckon right now the Enterprise concern is more FOMO around the AI wave and how to retrain or replace up to hundreds of thousands of employees. I don't think cost is the main concern right now. But if AI doesn't lead quickly to vast large scale replacement of workers as promised, I could definitely see the C-suits and their gaggle of consultants starting to ask questions about token pricing.
- LUmBULtERA 4mo agoAPI token price is one thing, but subscriptions on Claude are a good value. Weirdly everyone says that Claude subscriptions are subsidized because of the API price, even though (1) no one actually knows Claude's cost of inference, and (2) Chinese providers are also able to provide cheap inference, so why do they think Claude can't? I also wonder if Enterprises have deals for other API pricing that is not posted publicly, so all we see is a high API sticker price.
- wuliwong 4mo agoI only have knowledge of one enterprise deal but there is no discount. Which I found surprising.
- mnicky 4mo ago> no one actually knows Claude's cost of inference There were some rumors stating that their margin is around 70%. So they could go much cheaper probably, talking inference only. The other thing is R&D cost...
- gruez 4mo agoYour question relies on the premise that Chinese companies continue releasing free models. What's "the moat" for them continuing to do that?
- michaelcampbell 4mo ago> while being only marginally better. It's only marginally better in the things it's actually comparable to. A\ models are MUCH better in many more things; eg: things Kimi/etc. didn't distill. For those things the difference is like a cliff.
- tornikeo 4mo agoThat's a baseless claim that borderline reads like shilling. Do you have any proof of that you wrote there?
- bensyverson 4mo agoPart of Anthropic's moat is Claude Cowork & Claude Code. They got coders comfortable with CC and enterprise users comfortable with Cowork, and both are creating stickiness. The reality is that $20/$100/$200/mo feels reasonable to a lot of people relative to the value they're getting out of Claude, and if they switch to something else, there's a risk that it won't be as good, and they'll have a new tool to learn. It's not an insurmountable moat, but don't underestimate the user experience. The iPod didn't win because it was the cheapest device or the one with the most features.
- selfawareMammal 4mo agoPerformance. I pay for Opencode but none of the models give me Codex performance, so I have to keep my 20€ subscription+ the Opencode one
- bgins 4mo agoI am still very new to the open-weight/source models. If anyone is using them full-time, I’d really love to hear about the setup and how they perform, as I am considering moving my org off Anthropic products.
- trollbridge 4mo agoQwen 3.6 seems to be the strongest local models, works OK on an RTX 5090 or a > 32GB Mac.
- andai 4mo agoI keep trying to switch to the Chinese models, but I keep finding myself asking Claude to fix their outputs. (Both functionality and style.) So I always end up switching back.[0] I also keep trying GPT, which is quite solid. Very fast, great at debugging. But its code is often overly clever and hurts my brain. (Maybe fixable with prompting. I tried and it helped the Chinese ones a bit. Just tell them do be elegant, like in the old image AI days "+good -bad"!) For now I do still need my human brain to actually be able to make sense of the stuff, and Claude is the only one that consistently meets that requirement. But I am hoping that one of these days, one of the Chinese labs figures out the special sauce :) -- [0] (For smallish edits, though, I am having a great time with DeepSeek Flash. Practically unlimited AI on tap! How cool is that.)
- scottcha 4mo agoI use glm5.1 plus pi with a few customized skills and am very happy with it. I hadn’t touched my Claude 5x plan for a couple of weeks but opened it back up in Claude code when fable was released and did a few tasks and still was happy to return to glm/pi.
- sebastianconcpt 4mo agoBetter than Qwen3.6-35B-A3B-8bit ? When I tried glm found it way way slower (omlx as runtime)
- scottcha 4mo ago
- 343rwerfd 4mo agoI think any new model not demonstrably maybe 20-30% over Deepseek v4 capabilities priced over the price per token of Deepseek is almost automatically deprecated as low use model (maybe for Planning).
- giancarlostoro 4mo agoIs Deepseek just eating cost or are people able to host their open models for comparable costs?
- re-thc 4mo agoThey focused on caching and other optimizations.
- psittacus 4mo agoIf openrouter is to be trusted, the cheapest offers that are not from Deepseek itself are: - twice as expensive on the output (1.52 vs 0.87) - six times as expensive on the input (0.33 vs 0.05) https://openrouter.ai/deepseek/deepseek-v4-pro?sort=price#pricing https://openrouter.ai/deepseek/deepseek-v4-pro?sort=price#pr...
- trollbridge 4mo agoOther people are hosting it in the same order of magnitude. Xioami recently matched DeepSeek’s pricing.
- rsanek 4mo agoLikely CCP-subsidized
- natrys 4mo agoThese things enormously benefit from economies of scale. I am fairly certain their margins might be low but they don't actually sell API at loss, however that doesn't mean your cost footprint would be anywhere as low.
- 0xbadcafebee 4mo agoDeepSeek v4 Pro is not actually that good a model compared to GLM 5.1 and Kimi K2.6. It's an okay coder/thinker for the price.
- shreedx 4mo agoI would really love to know if anyone has any experience with something like opencode + Kimi K2.6/2.7 now compared to Claude Code. What is better, what is worse, what is the cost comparison. I am currently paying $100 for the 5x Max plan, but Fable is running through the usage limits quite drastically and I cannot really say it's night and day compared to Opus. Also, I use this mostly for my side projects, so the $100 bill is quite noticeable. I definitely don't want to pay more.
- ramon156 4mo agoI can only talk about GLM 5.1 which is roughly at sonnet 4 levels imo. It's good, does most tasks well that I throw at it, but will fail at anything congitive/complex. It gets stuck often. It costs ~6$ a month though
- jeremyjh 4mo agoThis was my experience using GLM 5.1 in Claude Code but it works far better in OpenCode, I’d really like to understand why. I think it’s a bit stronger than Sonnet 4.6. I use the oh-my-openagent planning system and haven’t used vanilla OpenCode enough to know how much that is contributing.
- miroljub 4mo agoThe answer is easy, CC is bug for bug optimized for Anthropic models. They don't even test it with other models, let alone provide support for all small compatibility quirks of different provider implementations. On the other hand, Opencode, Pi agent and other open source tool offer much better support for all models, including open source.
- re-thc 4mo agoThe Kimi problem is it doesn’t follow instructions and goes off track often. Other than that it’s pretty decent (for the price).
- 4mo ago
- jackdoe 4mo agoI think there is some threshold after which "best" model doesn't matter, we are not that far from it. Fable now is really good, in a year or so, if Kimi catches up, even if Fable6 is much better, I think I will use kimi at 1/10th of the price. I said that about opus 4.5 at the time, thinking "this is so good, in 6-12 months the Chinese models will be as good and cheap, I will use them", but I was wrong.. I pay premium for opus4.7/8 and Fable. But at some point, it will just do the thing you want it to do, and then the race to the bottom will start. Now that Chinese companies have access to some very good Fable tokens, I hope it speeds up the race.
- Zoadian 4mo agoprice/token isnt the only thing relevant. if you have to ask the AI again, it'll cost you more than when it gets things right in the first place. so better models may still be cheaper even if the price per token is higher.
- jackdoe 4mo agoyes, that is my point, but at some point, better is unmeasurable, and both the better and the not-as-good produce similar result, and then you pick the one with 1/10th of the price
- wolttam 4mo agoDepending on who you are and how you use these models, we're already at this point
- xendo 4mo agoExactly, for long running vibe coded stuff that I don't care about quality getting big and smart model is the only option. But for high quality changes where I need to have control and understand everything, where I do everything in small chunks - I can use basic model like Sonnet.
- deleted 4mo ago[deleted]
- haeseong 4mo ago[flagged]
- fractalf 4mo agoHow is 2.7 a thing _now_ ? it's not even mentioned on moonshot's webpage..
- cassianoleal 4mo agoIt's not 2.7. It's 2.7-Code, and it's 2.6 token-optimised for coding. https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstart https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstar...
- jkwang 4mo agoThis maps to what I'm seeing in practice. The gap between demo and production is consistently underestimated, especially around error handling and edge cases.
- goldenarm 4mo agoBenchmark geometric mean - GPT-5.5: 62.7% - Opus 4.8: 62.2% - Kimi K2.7 Code: 56.3% - Kimi K2.6: 48.2%
- lostmsu 4mo agoWould be nice to have 5.2 and 4.6 for comparison.
- RobertPelloni 4mo agoinsanely great!
- jdw64 4mo agoPersonally, when I use open code or routers, I feel that beyond a certain level, the models don't make a huge difference to me. Except for expensive and mediocre models like Gemini. In that sense, Chinese models are pretty good. I usually write code in function or method units and then design and assemble them together. GPT series models are more thorough and better, but I'm not sure if the difference is enormous. It seems to depend on the workflow, but in my opinion, if you are thorough enough, I wonder if there really is a big difference
- onlyrealcuzzo 4mo agoIn my experience, there's little difference between implementing individual functions between frontier models and SotA ~30B param models. Once you have a coherent design (the hard part), you can feed it to a pretty small model and get basically the same quality. They'll not one-shot, but they're faster and cheaper, so it still works out in your favor. Plus you can do it locally...
- jdw64 4mo agoI have a similar experience. However, when including code review, I think the GPT model is the most impressive
- dcreater 4mo agoI really hope we stop using the term "Chinese models". It has this air of Negative connotation. It's the equivalent of calling cars Japanese, which people used to do but now is almost entirely meaningless. You just call them Toyota, Honda, Lexus etc.
- jdw64 4mo agoYou are right. I agree.It may seem like a kind of bias, but I hadn't thought of that part. Thank you for pointing out my bias.
- theanonymousone 4mo ago
- giancarlostoro 4mo agoReading their modified license terms, it cracks me up, because they've basically remade the MIT to be the MIT + the one clause that the BSD used to have, which didn't care about MAU or revenue, if you used it in a product, they asked you to 'advertise' them basically. Honestly, its a reasonable request.
- htrp 4mo agoThis is the cursor callout. Don't make us shame you into disclosure
- giancarlostoro 4mo agoAh is that what it is? I don't use Cursor, never saw it as being relevant to me, but would not surprise me.
- schmorptron 4mo agoCursor's composer models are finetuned kimi
- varispeed 4mo agoThey are unusable (unless you want to deliberately destroy your codebase). So if Cursor's models are Kimi based, then well. I'll skip them altogether.
- jingpostmedia 4mo ago[flagged]
- jingpostmedia 4mo ago[flagged]
- RIshabh235 4mo agoI think deepseek has crossed the threshold for being on par with opus 4.6 and kimi is doing a great job in shipping velocity.
- pixel_popping 4mo agoDeepseek V4 is far from Opus 4.6 level, it might look like it at first glance, but the general reasoning (especially multi-steps) is frankly far off. It's good enough to build great things don't get me wrong, but there is really something that is different from Anthropic models.
- RIshabh235 4mo agoagreed
- deleted 4mo ago[deleted]
- minraws 4mo agoI tested it properly and it seems rather decent improvement atleast it does use less tokens for the same task which is good enough a reason for me to use it over k2.6 if I need an open model
- pcwelder 4mo agoGreat! Finally follows custom tool call format (k2.6 couldn't). It's a good indicator of instructions following and agentic behaviour. UIs it's generating is pretty good, not without problems, but certainly better than other models at this price point.
- Bolwin 4mo agoWhat do you mean by custom format? Non-json?
- pcwelder 4mo agoCould be json or non json. Instead of using tools in API, you ask model to share structured output in text. You parse the string to get the JSON. Gives much more control over things you can do. For example model shares <tool_call name="getWeather"> <param name="city">London</param> </tool_call>
- theanonymousone 4mo agoIn OpenRouter, there is an "int4" tag for Moonshot provider of Kimi K2. 7 Code. Isn't that too low, particularly coming from the very developer of the model? Os that a mistake? How is it in their direct API offer?
- kouteiheika 4mo agoThe model is natively quantized (i.e. it was trained that way in the first place, so this is not a post-training quantization which degrades performance).
- theanonymousone 4mo agoBut the huggingface link mentions BF16, F16, and I32?
- kouteiheika 4mo agoNot every weight is quantized. For example, those weights which don't take much space or are highly important are left in higher precision. State-of-art quantization of weights is never done uniformly (i.e. to all weights and in the same way).
- zackangelo 4mo agoI don't believe safetensors has a native int4 dtype, so they packed 4 int4s into a bf16 in this checkpoint.
- knollimar 4mo agoIsn't it not completely quantized? I thought there were some dense parts but most is int4?
- wgd 4mo agoOften in MoE models the experts are quantized while the shared portions, being a much smaller part of the network with greater impact, are kept at higher or full precision. Not familiar with the Kimi QAT approach specifically but it's likely they do this.
- Bnjoroge 4mo agoOutput tokens are almost 5x more expensive than mimov2.5 pro/dsv4pro. I’m curious to see if Kimik2.7 is that much better. Feels like kimi are positioning themselves as the premium open source models
- mdasen 4mo agoI find that I don't use a ton of output tokens. I'm usually around 95% cached input, 4% input, and 1% output. For me, the big thing with MiMo-V2.5-Pro and DeepSeek V4-Pro is that cached inputs are practically free. Kimi K2.7 Code is 53x more expensive for cached inputs which is 95% of my costs. If I use 95M cached input tokens, 4M input tokens, and 1M output tokens, that'd be: $18 for cached input on Kimi K2.7 Code vs $0.34 with MiMo/DS; $3.80 for inputs on Kimi vs $1.74 with MiMo/DS; and $4 for output on Kimi vs $0.87 with MiMo/DS. Of all the places where I'm accumulating costs by using Kimi, it's the cached inputs. The real savings with MiMo/DS's price cut is the cached inputs.
- wolttam 4mo ago95/4/1 holds here too
- btian 4mo agoIt's not more expensive at all. They are all open weights models. I run them on 2x8xH100. They cost the same.
- Bnjoroge 4mo agoOpenrouter has them as significantly more expensive.
- SubiculumCode 4mo agoHas anyone taken these open weight models from China and stripped the CCP out of them? I do not mean that snarkily, I mean review them thoroughly using techniques for weight introspection (concept activations) in response to things that one might expect would trigger deceptive/malicious behavior if the CCP had actually tried to implant context-specific behaviors (e.g. the accusation of generating vulnerable code if being used in American government applications, which I don't know if it was ever proven). Just in case there are those who'd reflexively down vote this post, I'd just like to say that in a time of great national geopolitical rivalries, this kind of question is not unreasonable one to ask. Indeed, its applicable question whichever nation you live in.
- justinclift 4mo agoSounds like something that heretic or similar might be useful for? https://github.com/p-e-w/heretic https://github.com/p-e-w/heretic
- threethirtytwo 4mo agoEh even corporate created LLMs are suspect to corporate biases. Nothing is safe.
- SubiculumCode 4mo agoEverything is the same is not a serious argument because they are not the same.
- threethirtytwo 4mo agoThey are different and yet the same. The biggest difference is there’s generally more hatred for China because many us citizens are jealous. But corporate corruption is not that different in safety. Other than hatred the difference lies in incentives. Corporations want profit. China just wants to spy.
- SubiculumCode 4mo ago
- storus 4mo agoIs this Moonshot.ai's attempt to replicate Composer 2.5 (coding fine-tune of Kimi 2.5) from Cursor IDE?
- Symmetry 4mo agoI wish they wouldn't call these "open source" models. The output weights are open but that's more analogous to a binary. The source would be the training data and techniques that went into producing the binary/weights. "Open weights" is also a term in wide use and accurately tells us what we're getting.
- Eridrus 4mo agoIt's not quite as closed as a binary, it is very standard practice to take these models and fine-tune them. If there were actually even close to frontier open source models, this would be more of a discussion, but everyone knows these mean open weight.
- pizlonator 4mo agoI just had Kimi K2.7-code rebase my Fil-C OpenSSL patch from 3.3.1 to 3.5.7 with quite bare bones instructions and it seems to have worked. 177KB patch, so it's not a small change. The patch did not apply cleanly initially; the agent had to do nontrivial work. I just showed it the patch against 3.3.1, what command to use to build, and the path to 3.5.7 along with a link to the documentation of the change (https://fil-c.org/constant_time_crypto https://fil-c.org/constant_time_crypto). Note, I use my own coding agent (T800, which isn't public, and was previously well tested and tuned for K2.5). I think this cost me between $5 and $10 in API usage. (EDIT: OpenSSL, not OpenSSH)
- tomaytotomato 4mo ago"T800" Do you have your agent say things like "Hasta la vista baby", or "I'll be back, after I clear my context" ?
- pizlonator 4mo agoYes
- XCSme 4mo agoSeems to be similar level to Kimi K.26, just that it's more token efficient and cheaper to run: https://aibenchy.com/compare/moonshotai-kimi-k2-6-medium/moonshotai-kimi-k2-7-code-medium/ https://aibenchy.com/compare/moonshotai-kimi-k2-6-medium/moo...
- madduci 4mo agoLooks interesting but yet no Ollama model?
- deleted 4mo ago[deleted]
- zftnb666 4mo ago[flagged]