15 ms·
Kimi K2.6: Advancing open-source coding
- irthomasthomas 6mo agoBeats opus 4.6! They missed claiming the frontier by a few days.
- NitpickLawyer 6mo agoWhile I'm skeptical of any "beats opus" claims (many were said, none turned out to be true), I still think it's insane that we can now run close-to-SotA models locally on ~100k worth of hardware, for a small team, and be 100% sure that the data stays local. Should be a no-brainer for teams that work in areas where privacy matters.
- cedws 6mo agoEven the smaller quantized models which can run on consumer hardware pack in an almost unfathomable amount of knowledge. I don't think I expected to be able to run a 'local Google' in my lifetime before the LLM boom.
- sterlind 6mo agoI'm extremely curious how these models learn to pack a lossily-compressed representation of the entire Internet (more or less) into a few hundred billion parameters. like, what's the ontology?
- osti 6mo agoI think this one is only about 600GB VRAM usage, so it could fit on two mac studios with 512GB vram each. That would have costed (albeit no longer available) something like less than 20k.
- NitpickLawyer 6mo agoYeah, but that's personal use at best, not much agentic anything happening on that hardware. Macs are great for small models at small-medium context lengths, but at > 64k (something very common with agentic usage) it struggles and slows down a lot. The ~100k hardware is suitable for multi-user, small team usage. That's what you'd use for actual work in reasonable timeframes. For personal use, sure macs could work.
- osti 6mo agoTrue, but I think for local models, we are mostly considering personal usage.
- zozbot234 6mo agoYou could run it with SSD offload, earlier experiments with Kimi 2.5 on M5 hardware had it running at 2 tok/s. K2.6 has a similar amount of total and active parameters.
- osti 6mo agoYeah... I would definitely call 2t/s unusable. For simple chats, I'd want at least 15 t/s. For agentic coding (which this model is advertised for), I'd want good prefill performance as well.
- veber-alex 6mo agoThat's just throwing money away. The performance with large context would have been unusable especially if you need to serve more then a single person.
- BoorishBears 6mo agoOpus is clearly a sidegrade meant to help Anthropic manage cost, so I would say they may have it if it actually beats 4.6
- irthomasthomas 6mo agoCould be right. I just noticed my feed is absent the usual flood of posts demoing the new hotness on 3D modeling, game design and SVG drawings of animals on vehicles.
- pixel_popping 6mo agoIt doesn't beat Opus 4.6, no way, don't be fooled by benchmarks.
- nickandbro 6mo agoWow, if the benchmarks checkout with the vibes, this could almost be like a Deepseek moment with Chinese AI now being neck and neck with SOTA US lab made models
- motoboi 6mo agoWith the previous generation? Yes. With 10T mythos-level models? Not even close.
- bestouff 6mo agoThere's no public data about Mytho.
- maplethorpe 6mo agoThat's because it would be too dangerous to release.
- nisegami 6mo agoSo is my P=NP proof.
- cedws 6mo agoMy girlfriend goes to a different school, you wouldn't know her.
- squarefoot 6mo agoSame for teleport, time travel and warp drive.
- rockinghigh 6mo agoThey could release data to back up that claim.
- amazingamazing 6mo agoThe psyop continues. Mythos until it’s released is vaporware. Notice how you can try kimi 2.6. Where is the same for mythos?
- swingboy 6mo agoExciting benchmarks if true. What kind of hardware do they typically run these benchmarks on? Apologies if my terminology is off, but I assume they're using an unquantized version that wouldn't run on even the beefiest MacBook?
- esafak 6mo agoK2.5 was already pretty decent so I would try this. Starting at $15/month: https://www.kimi.com/membership/pricing https://www.kimi.com/membership/pricing edit: Note that you can run it yourself with sufficient resources (e.g., companies), or access it from other providers too: https://openrouter.ai/moonshotai/kimi-k2.6/providers https://openrouter.ai/moonshotai/kimi-k2.6/providers
- wg0 6mo agoHow are the usage limits compared to Anthropic?
- greenavocado 6mo agoAnthropic has the worst usage limits in the industry
- andriy_koval 6mo agogemini is worse imo
- deaux 6mo agoYou're correct, Gemini chat limits are a joke at their chapest paid tier compared to both Claude and GPT. Especially crazy when you consider Gemini 3 Pro is more than twice as cheap as Opus 4.6 on the API. It's hard to run into pure chat limits on Claude even if you only use Opus on the cheapest tier, whereas with Gemini it's easy to hit. Not sure about coding usage, Google being weird about these things I could see that quota being separate.
- gessha 6mo agoI’m not sure what A/B test you’re part of but on Claude Code Pro, I hit every single one of my quotas without exception. If you analyze/process images it’s even worse: I hit rate limits first and if I use separate sessions, I hit my quotas too. I use up so many tokens that Jensen should hire me.
- deleted 6mo ago[deleted]
- lbreakjai 6mo agoI have a subscription through work, I've been trialing it, so far it looks on par, if not better, than opus.
- verdverm 6mo agohttps://huggingface.co/moonshotai/Kimi-K2.6 https://huggingface.co/moonshotai/Kimi-K2.6 Is this the same model? Unsloth quants: https://huggingface.co/unsloth/Kimi-K2.6-GGUF https://huggingface.co/unsloth/Kimi-K2.6-GGUF (work in progress, no gguf files yet, header message saying as much)
- Balinares 6mo agoQuite curious how well real usage will back the benchmarks, because even if it's only Opus ballpark, open weights Opus ballpark is seismic.
- gpm 6mo agoHuh, so the metadata says 1.1 trillion parameters, each 32 or 16 bits. But the files are only roughly 640GB in size (~10GB * 64 files, slightly less in fact). Shouldn't they be closer to 2.2TB?
- johndough 6mo agoThe bulk of Kimi-K2.6's parameters are stored with 4 bits per weight, not 16 or 32. There are a few parameters that are stored with higher precision, but they make up only a fraction of the total parameters.
- gpm 6mo agoHuh, cool. I guess that makes a lot of sense with all the success the quantization people have been having. So am I misunderstanding "Tensor type F32 · I32 · BF16" or is it just tagged wrong?
- rockinghigh 6mo agoThe MoE experts are quantized to int4, all other weights like the shared expert weights are excluded from quantization and use bf16.
- liuliu 6mo ago
- pt9567 6mo agowow - $0.95 input/$4 output. If its anywhere near opus 4.6 that's incredible.
- corlinp 6mo agoThis should erase any doubt that AI Labs are making $$$ on API inference. Kimi 2.5 (which this is based on) is served at $0.44 input / $2 output by a ton of different providers on OpenRouter, 2.6 will certainly be similar. That's about 11X less than Opus for similar smarts.
- Lalabadie 6mo agoFamously, OpenAI and Anthropic are devoted to increasing efficiency before scaling up resource usage.
- amazingamazing 6mo agoHow does it erase any doubt? You’re implying Chinese things can’t be actually cheaper to produce than American which is laughable
- corlinp 6mo agoMost of those inference providers are American, and China is actually at a disadvantage here because of export restrictions - US companies are using newer and more efficient chips.
- amazingamazing 6mo agoIf it’s newer and efficient then why is the api more expensive?
- veber-alex 6mo agoPrice is set based on what people are willing to pay not based on actual costs.
- greenavocado 6mo agoI pray the benchmark figures are true so I can stop paying Anthropic after screwing me over this quarter by dumbing down their models, making usage quotas ridiculously small, and demanding KYC paperwork.
- jollymonATX 6mo agoAnthropic has done horrible PR and investors should be livid.
- greenavocado 6mo agoMy theory is they pushed retail off their systems to make room for their new corporate fat cat clients. In which case, they'll do just fine.
- deaux 6mo ago> dumbing down their models, This should be so easy to prove if it were true. Yet there is none of it, just vibes. Still, your other two points are completely valid. The opaqueness of usage quotas is a scam, within a single month for a single model it can differ by more than 2x. And this indeed has been proven.
- greenavocado 6mo ago> This should be so easy to prove if it were true. https://github.com/anthropics/claude-code/issues/42796 https://github.com/anthropics/claude-code/issues/42796 https://scortier.substack.com/p/claude-code-drama-6852-sessions-prove https://scortier.substack.com/p/claude-code-drama-6852-sessi...
- deaux 6mo agoFirst link is about the harness, Claude Code, defaulting to less thinking over time. This isn't "the model getting worse". Second link is just a discussion of the first link.
- 6mo ago
- elfbargpt 6mo agoI've always been surprised Kimi doesn't get more attention than it does. It's always stood out to me in terms of creativity, quality... has been my favorite model for awhile (but I'm far from an authority)
- regularfry 6mo agoDirt cheap on openrouter for how good it is, too. Really hoping that 2.6 carries on that tradition.
- varispeed 6mo agoMaybe because it's a bit of like unleashing a chaos monkey on your codebase? I tried it locally (K2.5 72B) and couldn't get anything useful.
- KaoruAoiShiho 6mo agoHuh, that's not a thing?
- johndough 6mo agoThe parent poster is probably referring to Kimi-Dev-72B¹, which is a much smaller and older model, while people are probably more familiar with the big and fairly powerful 1100B Kimi-K2.5². [1] https://huggingface.co/moonshotai/Kimi-Dev-72B https://huggingface.co/moonshotai/Kimi-Dev-72B [2] https://huggingface.co/moonshotai/Kimi-K2.5 https://huggingface.co/moonshotai/Kimi-K2.5
- natrys 6mo agoYes it was good for its time, but 10 months old now which is a long time ago in this space. It was also a fine-tune (albeit a good one) of Qwen-2.5 72B. I wish they did more smaller models. Kimi Linear doesn't really count, it was more of a proof of concept thing.
- culi 6mo ago
- game_the0ry 6mo agoThere is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.
- osti 6mo agoMaybe open source == communism
- darkwater 6mo agoGood ol' Steve "Developers! Developers! Developers!" Ballmer said so a long time ago. What a visionary!
- konart 6mo agoBut China is not communist event though the rulling party the word in its name.
- osti 6mo agoOh i’m fully aware of that lol
- fragmede 6mo agoThe Democratic People's Republic of Korea would like a word.
- pheggs 6mo agowhat makes you think that china ever gave up its communist goals? I personally see that everything they do aims towards that goal. From the one child policy, the huge amounts of empty apartments they build, the stuff they produce for almost free, the fishing.. open sourcing the models perfectly fits that culture too, it's the means of production
- otterley 6mo agoThe one-child policy died a long time ago. Also, the accumulation of wealth by connected politicians and businesspeople flies in the face of what communism is supposed to stand for. There is a reason real estate values in popular cities has skyrocketed, and it’s not due to the locals getting wealthier. It’s where Chinese and other oligarchs put their ill-gotten wealth (well, besides Bitcoin).
- nisegami 6mo agoThe choice of example task for Long-Horizon Coding is a bit spooky if you squint, since it's nearing the territory of LLMs improving themselves.
- Banditoz 6mo agoIf the benchmarks are private, how do we reproduce the results? I looked up the Humanity's Last Exam (https://agi.safe.ai/ https://agi.safe.ai/) this model uses and I can't seem to access it.
- johndough 6mo agoYou can request access here: https://huggingface.co/datasets/cais/hle https://huggingface.co/datasets/cais/hle The test data is purposely difficult to access to reduce the chance of leaking it into the training dataset.
- mariopt 6mo agoReally excited to try this one, I've been using kimi 2.5 for design and it's really good but borderline useless on backend/advanced tasks. Also discovered that using OpenCode instead of the kimi cli, really hurts the model performance (2.5).
- oliver236 6mo agoisnt this better than qwen?
- Alifatisk 6mo agoWe'll have to wait for the results on Artificial analysis
- simonw 6mo agoAccessed via OpenRouter, this one decided to wrap the SVG pelican in HTML with controls for the animation speed: https://gisthost.github.io/?ecaad98efe0f747e27bc0e0ebc669e94/pelican.html https://gisthost.github.io/?ecaad98efe0f747e27bc0e0ebc669e94... Transcript and HTML here: https://gist.github.com/simonw/ecaad98efe0f747e27bc0e0ebc669e94#2026-04-20t164936----conversation-01kpnwt8d2bt5qwkm60j9sbkbs-id-01kpnwra0prz6v822cct5b08kq https://gist.github.com/simonw/ecaad98efe0f747e27bc0e0ebc669...
- SwellJoe 6mo agoWe got an overachiever, here. Kimi sounds like a teacher's pet kind of name.
- subscribed 6mo agoUnderappreciated comment
- FlyingSnake 6mo agoAt this point drawing these Pelicans must be in the training data sets.
- ffsm8 6mo agoClearly not. I mean the prompt was succinct and clear, as always - and it still decided to hallucinate multiple features (animation + controls) beyond the prompt. It'd also like to point out that to date no drawing was actually good from an actual quality perspective (as in comparative to what a decent designer would throw together) Theyre always only "good" from the perspective of it being a one shot low effort prompt. Very little content for training purposes.
- nwienert 6mo agoThe way I’ve come to think of LLM is that what the produce in a single reply even with thinking turned up, is akin to what you’d do in a single short session of work. And so if you ask it to do something big it will do a very surface level implementation. But if you have it iterate many times, or give it small pieces each time, you’ll end up with something closer to what a human would do. I imagine the pelican test but done in a harness that has the agents iterate 10+ times would be closer to what you’d expect, especially if a visual model was critiquing each time.
- dmix 6mo agoI'm pretty Kimi is what Cursor uses for their "composer 2" model. Works pretty good as a fallback when Claude runs out, but definitely a downgrade.
- arcanemachiner 6mo agoIt's a Kimi K2.5 finetune, there was some drama about this a few weeks ago.
- dmix 6mo agoWhat was the drama about?
- arcanemachiner 6mo agoThey were not open about the fact that it was trained on Kimi K2.5. This link explains it better than I can: https://www.trendingtopics.eu/cursor-admits-composer-2-is-built-on-chinese-ai-model-kimi-k2-5/ https://www.trendingtopics.eu/cursor-admits-composer-2-is-bu...
- 59nadir 6mo agoCursor seemingly went out of their way to not mention that they were actually running Kimi K2.5 and essentially by that omission made it seem like they had made their own model. They added a note to a blog post about using it at some point and then when they wrote a new one they conveniently left it out again. That's at least what I perceived as "the drama".
- deleted 6mo ago[deleted]
- cassianoleal 6mo agoIf only their API wasn't tied to a Google or phone login...
- jenkstom 6mo agoIf it's open then there will be multiple providers. I see it is on OpenRouter now.
- cassianoleal 6mo agoI'm going to experiment with this, but unless it's insanely more efficient in token usage than anything else I've tried, the only way to keep costs more or less acceptable is through a subscription.
- atemerev 6mo agoWhy use "their API"? It is an open model, use any provider on OpenRouter
- wolttam 6mo agoBecause sometimes (a lot of the time in my experience) third-party providers and inference engines fail to implement the model correctly in ways that are sometimes very subtle and not obvious. Deepinfra for example is not preserving thinking correctly for GLM5.1, even though they are for GLM5. This is one of the more obvious issues that crop up.
- polski-g 6mo agoYeah some of the providers on Openrouter correctly list what quantization they offer. And some refuse to say. OR should kick them off their platform if they want to be secretive.
- cmrdporcupine 6mo agoRunning it through opencode to their API and... it definitely seems like it's "overthinking" -- watching the thought process, it's been going for pages and pages and pages diagnosing and "thinking" things through... without doing anything. Sitting at 50k+ output tokens used now just going in thought circles, complete analysis paralysis. Might be a configuration or prompt issue. I guess I'll wait and see, but I can't get use out of this now.
- jbaiter 6mo agoHad the same experience using it for a refactor of a 3k LOC monolith via the Pi harness and OpenRouter. After burning through $8 worth of tokens it left the code in a broken state, the "thoughts" were full of loops where it would edit the monolith, then refer back to the original file, not finding it and then overwriting its changes with "git checkout --"
- cmrdporcupine 6mo agoIt's probably bad harness. I had a similar bad experience with qwen max yesterday also through opencode. In the past I tried Kimi thru Claude code I might try that again
- sankalpmukim 6mo agoI think this kind of overthinking is an extremely common pattern in the Chinese models. GLM's models are also very much like this.
- m4rkuskk 6mo agoI have been testing it in my app all morning, and the results line up with 4.6 Sonnet. This is just a "vibe" feeling with no real testing. I'm glad we have some real competition to the "frontier" models.
- mchusma 6mo agoit feels like between K2.6 and GLM5.1 we have Sonnet level intelligence at roughly Haiku level pricing. Which is great. I'm hoping that Anthropic will be able to release an updated Haiku soon and they really need something that is 1/3-1/5 the price of Haiku to compete with the truly cheaper models (Gemma-4 is really good at this range).
- XCSme 6mo ago(commented on the wrong thread, HN doesn't let me delete it :( )
- wizee 6mo agoThey're comparing to Opus 4.6, not 4.5. It was Anthropic's best public model up until last week.
- candl 6mo agoAre there any coding plans for this? (aka no token limit, just api call limit). Recently my account failed to be billed for GLM on z.ai and my subscription expired because of this... the pricing for GLM went through the roof in recent months, though...
- wolttam 6mo agoKimi has their own subscription that works basically the same as all the others. https://www.kimi.com/code https://www.kimi.com/code
- fg137 6mo agoAt $19/month, hard to see why I want to use Kimi over Claude.
- randomtoast 6mo agoBecause Opus on $20 CC is a joke. The $19 plan on Kimi has actually workable usage limits.
- plutokras 6mo agoThey tick all the boxes I care about – desktop & mobile app, cli – but for the same price I might as well just go for the leading providers.
- cute_boi 6mo agofor similar plan i think claude costs like $100 a month?
- phainopepla2 6mo agoClaude usage at $20 is basically unusable for serious work. I haven't used Kimi but I'd have to imagine they're offering a good deal more usage for the same price.
- ankit70 6mo agoYou can use $20 pro plan on Ollama or $10 one on OpenCode Go. Both has Kimi 2.6 live. https://opencode.ai/go https://opencode.ai/go https://ollama.com/pricing https://ollama.com/pricing
- kburman 6mo agoHas anyone here used Kimi for actual work? I tried it once, although it looks amazing on benchmarks, my experience was just okay-ish. On the other hand, Qwen 3.6 is really good. It’s still not close to Opus, but it’s easily on par with Sonnet.
- deanc 6mo agoYes. You’re using Kimi if you use the composer-2 model in cursor. It’s great. Plan in state of the art. Execute in composer-2
- rubslopes 6mo agoBefore GLM-5.1, I was going back and forth between Opus 4.5 and Kimi 4.5 and having very good results with Kimi.
- try-working 6mo agoI've used Kimi K2.5 when I run out of Codex quota. It does small and medium things OK. But if I work on complex things, I'll later have to spend two days cleaning up the mess with Codex. Hopefully 2.6 does better.
- antirez 6mo agoHere I analyze the same linenoise PR with Kimi K2.6, Opus, GPT. https://www.youtube.com/watch?v=pJ11diFOjqo https://www.youtube.com/watch?v=pJ11diFOjqo Unfortunately the generation of the English audio track is work in progress and takes a few hours, but the subtitles can already be translated from Italian to English. TLDR: It works well for the use case I tested it against. Will do more testing in the future.
- dygd 6mo ago> Agent Swarms, Elevated: Match 100 Jobs and Generate 100 Tailored Resumes Model seems quite capable, but this use-case is just yikes. As if interviewing isn't already a hellscape.
- XCSme 6mo agoIn my tests[0] it does only slightly better than Kimi K2.5. Kimi K2.6 seems to struggle most with puzzle/domain-specific and trick-style exactness tasks, where it shows frequent instruction misses and wrong-answer failures. It is probably a great coding model, but a bit less intelligent overall than SOTAs [0]: https://aibenchy.com/compare/moonshotai-kimi-k2-6-medium/moonshotai-kimi-k2-5-medium/z-ai-glm-5-medium/anthropic-claude-opus-4-7-medium/ https://aibenchy.com/compare/moonshotai-kimi-k2-6-medium/moo...
- deepsquirrelnet 6mo agoI tried it on openrouter and set max tokens to 8192, and every response is truncated, even in non-thinking mode. Maybe there's an issue with the deployment, but in your link also shows it generates tons of output tokens.
- XCSme 6mo agoOh yeah, I just noticed, like 3x the reasoning tokens.
- jauntywundrkind 6mo agoI really wish some of these very-long-horizon runs were themselves open sourced (open released open access). Have the harness setup to do git committing automatically of the transcript and code, offload the git commit message making. Release it all. This sounds so so so cool. It would be so amazing to see this unfurl: > Kimi K2.6 successfully downloaded and deployed the Qwen3.5-0.8B model locally on a Mac. By implementing and optimizing model inference in Zig—a highly niche programming language—it demonstrated exceptional out-of-distribution generalization. Across 4,000+ tool calls, over 12 hours of continuous execution, and 14 iterations, Kimi K2.6 dramatically improved throughput from ~15 to ~193 tokens/sec, ultimately achieving speeds ~20% faster than LM Studio.
- throwaw12 6mo agoBeats Opus and Open Source? I really hope this holds true in real world use cases as well and not only benchmarks. Congrats to Kimi team!
- Topfi 6mo agoK2.6-code-preview was a minor, but noticeable jump, especially in a long running testing task and prior Moonshot releases have been the only models that I'd consider a suitably competitive replacement for Anthropic models. The way they approach tool calls, task inference and adherence is far closer than any other providers output, similar to how GLM models map far more closely to OpenAIs releases. Whether task adherence, task assessment, task evaluation or task inference, K2.5 got closer to Opus 4.5 than any other model (but was still behind overall). I will have to test this full release of K2.6 but could see it serve as a very good overall drop-in replacement for Opus 4.5 and Opus 4.6 at 200k across the vast majority of tasks. I will say however that Opus 4.7 Max 1M has been a very significant jump in performance for me, especially in tasks beyond 120k token where I'd argue it is now the most reliable model in continued task adherence and tool calling without compaction. Ironically, my initial experience was less than pleasant as on XHigh I found task adherence to have regressed even with less than 1/10th of the context window having been used. Am very interested in K2.6s compaction strategy (which appears to be very simply all things considered) and how it performs beyond 100k tokens. As it stands, only OpenAI models have made compaction for long running tasks work well, though overall, GPT-5.4 is still inferior in my tests regardless of context window over other models such as Opus 4.6 1m and Opus 4.7 1m. Haven't gotten around to testing Opus 4.7 200k and will have to do this to properly assess K2.6 fairly, but I'd be very surprised if K2.6 truly beat Opus 4.7 200k given the jump I have experienced.
- Alifatisk 6mo agoDamn it, they stopped offering Kimmmmy. Their sales ai agent which allowed you to bargain for lower subscription prices.
- ttul 6mo agoAm I being paranoid in questioning whether the CPC would have something to gain by monitoring coding sessions with Chinese coding AI models? Coding models receive snippets of our intellectual property all day long. It's a bit of a gold mine, no?
- throwaw12 6mo agoI think you should worry more about NSA, FBI, ICE and other 3 letter US agencies monitoring your sessions
- ttul 6mo agoThere's nothing anyone can do about state-level espionage anywhere, using any cloud-hosted service. That being said, there is a very big difference between the legal situation in the United States vs. China. Chinese internet companies are required to have CPC interaction and since the rule of law does not strictly exist in China, the state can compel surveillance cooperation regardless of what might be written down. If a three-letter agency is compelling Anthropic to open up its queries for inspection, that kind of surveillance would be authorized by law and if Anthropic violated the law in cooperating, they would suffer the consequences in civil court. Maybe not immediately, but at least the possibility exists. In China, there's no recourse at all. Surveillance must be presumed.
- LordDragonfang 6mo ago> the rule of law does not strictly exist in China, the state can compel surveillance cooperation regardless of what might be written down While I agree that China is obviously worse in this regard, it's naive to claim this is unique to China, when literally a couple of months ago the US got into a fight with Anthropic about them not removing safeguards which were already just enforcing the letter of the law.
- tw1984 6mo agoRule of law in the US - are you kidding yourself? When American citizens are being gunned down in public on cameras by US federal government agents, you are telling me that the US follows the rule of law? Before you start to offer more propaganda, just tell me where is the killer of Renée Good, has that killer been arrested or charged yet? Keep your censored version of rule of law to yourself and your kids. oh, btw, the current US President did got convicted for criminal offences, he walked away for free just because he got elected as the president. nice rule of law! what did he do recently - authorised illegal war against another country in which over 100+ school children got killed. Surely your fancy US rule of law is going to do something about this?
- sixhobbits 6mo agoI tried it out with my normal mixed-up wolf, goat, cabbage problem and it couldn't solve it. Sonnet 4.6 also can't, but Opus 4.7 has no problems. Details here [0] [0] https://techstackups.com/comparisons/kimi-2.6-vs-opus-4.7-and-cabbages/ https://techstackups.com/comparisons/kimi-2.6-vs-opus-4.7-an...
- ninjahawk1 6mo agoI often wonder if in the future, the same way early computers used to take up an entire room but now fit in your pocket, if in the future the equivalent of a data center will be a single physical device like a phone nowadays. And if that’s the case, would it happen much quicker since technology has been speeding up year by year?
- gpm 6mo ago> And if that’s the case, would it happen much quicker since technology has been speeding up year by year? I wouldn't expect this. Historically we've had a roughly exponential rate of shrinkage. If we keep that same exponential going, we should expect the amount of time to shrink "room full of compute" to "pocket full of compute" to be equal. And recently we've fallen behind that exponential rate of shrinkage. And this is rather expected because exponentials are basically never sustainable rates of growth. I still expect that technological progress is getting faster year by year, and that we're still shrinking compute, but that's not necessarily enough for the next shrinking to take less time than when we had exponential progress on shrinking.
- Flux159 6mo agoThere’s some early work being done here by companies looking at making LLM ASICS like Taalas (HC1 gets 17k t/s for llama 8b - currently at 2.5kW which is closer to a single server, but this is their first chip). There’s other options like photonic computing which might be able to reduce power significantly but are still in research as far as I can tell. Because so much money is invested in AI & traditional gpu inference is so power hungry, I would expect significant improvements in this space quickly.
- OsamaJaber 6mo agoThe modified MIT clause is sneakier than people think. Hit 100M users or $20M a month and you have to slap "Kimi K2.6" on your UI. That covers any consumer app worth building. Not really open, more like free until you matter. Llama pulled the same move
- svachalek 6mo agoWorth building with VC capital maybe. A small team putting together an app that pulled in $20M per year should be pretty pleased with that.
- brightball 6mo agoThe threshold for "worth building" is much lower than that for a lot of people.
- codemog 6mo agoAnd the Kimi team broke the Anthropic ToS by training off Opus outputs and… nothing happened?
- darksaints 6mo agoNobody cares, nor should they. Anthropic broke nearly every ToS of every website that they scraped data from. The AI robber barons just want to monopolize intellectual property violations, and I'm gonna cheer on any robin hoods that take it back from them.
- dcchambers 6mo agoI'll definitely put this into the "good problem to have" category.
- Saline9515 6mo agoAttribution is a fair clause in opensource. What is the problem? You are making 20M$ a month thanks to their free work.
- throwaw12 6mo agoif you reach that numbers, kimi would be your least of worries
- dogscatstrees 6mo agoThis kimi website, it looks like a stylesheet from the 90's. They could learn a thing or two about typeface design. Steve Jobs would be incensed at this.
- kristianp 6mo agoI prefer a website that has the first page of text visible almost immediately, with no glitches when fonts load, tbh.
- deleted 6mo ago[deleted]
- rane 6mo agoAdded support for Kimi in https://github.com/raine/claude-code-proxy https://github.com/raine/claude-code-proxy and it does appear to work surprisingly well with Claude Code, although the usage limit for the entry tier doesn't seem as generous as I'd have expected.
- thomasahle 6mo agoDoes it run on Nvidia or Huawei?
- gertlabs 6mo agoEarly benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic workflows, but GLM 5.1 still handled them very well, so I'm curious how Kimi ends up. Both the old Kimi and the new model are on the slower side, so that's a consideration that makes them probably less usable for agentic coding work, regardless. The old Kimi K2 model was severely benchmaxxed, and was only really interesting in the context of generating more variation and temperature, not for solving hard problems. The new one is a much stronger generalist. Overall, the field of open weights models is looking fantastic. A new near-frontier release every week, it seems. Comprehensive, difficult to game benchmarks at https://gertlabs.com/?mode=oneshot_coding https://gertlabs.com/?mode=oneshot_coding
- cmrdporcupine 6mo agoSurprised to see such variance per language
- gertlabs 6mo agoIt's interesting; I can only speculate as to the underlying reason. When given enough time, models outperform in Rust/C++ in longer agentic tasks, and actually perform worst in Python. For tasks that aren't judged on code speed. https://gertlabs.com/?mode=agentic_coding https://gertlabs.com/?mode=agentic_coding
- edude03 6mo agoIt makes sense when you consider LLMs don't generalize very well, so they're heavily dependent on how good (how varied as well as how high quality) the training data is
- cmrdporcupine 6mo ago
- deleted 6mo ago[deleted]
- max2026 6mo ago[dead]
- waynevdm 6mo agoWith agents running at the scale and for an extended period. Surely they would need to pay for external services like APIs, compute, data. Would everything be based off subscriptions or API usage?
- potter098 6mo ago[flagged]
- v2space 6mo ago[dead]