11 ms·
Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
- billyellow 8mo agoCool
- mangolie 8mo agothey cooked
- jumploops 8mo ago> For complex tasks, Kimi K2.5 can self-direct an agent swarm with up to 100 sub-agents, executing parallel workflows across up to 1,500 tool calls. > K2.5 Agent Swarm improves performance on complex tasks through parallel, specialized execution [..] leads to an 80% reduction in end-to-end runtime Not just RL on tool calling, but RL on agent orchestration, neat!
- deleted 8mo ago[deleted]
- mohsen1 8mo agoParallel agents are such a simple, yet powerful hack. Using it in Claude Code with TeammateTool and getting lots of good results!
- esperent 8mo ago> TeammateTool What is this?
- jlu 8mo agoclaude code hidden feaure currently under a feature flag: https://github.com/mikekelly/claude-sneakpeek https://github.com/mikekelly/claude-sneakpeek
- frimmy 8mo agohttps://x.com/kieranklaassen/status/2014830266515382693 https://x.com/kieranklaassen/status/2014830266515382693 - agent swarms tool shipping w/ cc soon..
- XCSme 8mo ago> Kimi K2.5 can self-direct an agent swarm Is this within the model? Or within the IDE/service that runs the model? Because tool calling is mostly just the agent outputting "call tool X", and the IDE does it and returns the data back to AI's context
- mzl 8mo agoAn LLM model only outputs tokens, so this could be seen as an extension of tool calling where it has trained on the knowledge and use-cases for "tool-calling" itself as a sub-agent.
- XCSme 8mo agoOk, so agent swarm = tool calling where the tool is a LLM call and the argument is the prompt
- dcre 8mo agoSort of. It’s not necessarily a single call. In the general case it would be spinning up a long-running agent with various kinds of configuration — prompts, but also coding environment and which tools are available to it — like subagents in Claude Code.
- IanCal 8mo agoYes largely, although they’ve trained a model specifically for this task rather than using the base model and a bit of prompting.
- storystarling 8mo ago1,500 tool calls per task sounds like a nightmare for unit economics though. I've been optimizing my own agent workflows and even a few dozen steps makes it hard to keep margins positive, so I'm not sure how this is viable for anyone not burning VC cash.
- zozbot234 8mo ago"tool call" is just a reference to any elementary interaction with the outside system. It's not calling third-party APIs or anything like that.
- storystarling 8mo agoTrue, but that's still 1,500 inference cycles. Even without external API fees, the latency and compute burden seems huge. I don't see how the economics work there without significant subsidies.
- darrinm 8mo agoFWIW many tool calls can be and often are made in one inference cycle.
- DeathArrow 8mo agoThose are some impressive benchmark results. I wonder how well it does in real life. Maybe we can get away with something cheaper than Claude for coding.
- oneneptune 8mo agoI'm curious about the "cheaper" claim -- I checked Kimi pricing, and it's a $200/mo subscription too?
- mrklol 8mo agoThey also have a $20 and $40 tier.
- Alifatisk 8mo agoIf you bargain with their bot Kimmmmy (not joking), you can even get lower pricing.
- mohsen1 8mo agotell me more...
- Alifatisk 8mo agoGo to kimi chat, there will come up multiple suggestions of use cases. One of them will be the bargain robot. If you download their mobile app, the challenge to bargain will probably popup too! Depending on how well you bargain with the robot, you can go as low as 0,99$ (difficult). Either way, their moderate plan doesn’t have to be 20$. The agent wants a good reason for why it should lower the price for you. Here’s the direct link to Kimmmmy: https://www.kimi.com/kimiplus/sale https://www.kimi.com/kimiplus/sale I’ll send an invite link too if you don’t mind: https://www.kimi.com/kimiplus/sale?activity_enter_method=h5_share&invitation_code=2EVFCK&sharetype=link https://www.kimi.com/kimiplus/sale?activity_enter_method=h5_...
- spaceman_2020 8mo agoKimi was already one of the best writing models. Excited to try this one out
- Alifatisk 8mo agoTo me, Kimi has been the best with writing and conversing, its way more human like!
- Tepix 8mo agoHuggingface Link: https://huggingface.co/moonshotai/Kimi-K2.5 https://huggingface.co/moonshotai/Kimi-K2.5 1T parameters, 32b active parameters. License: MIT with the following modification: Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2.5" on the user interface of such product or service.
- Imustaskforhelp 8mo agoHey have they open sourced all Kimi k2.5 (thinking,instruct,agent,agent swarm [beta])? Because I feel like they mentioned that agent swarm is available their api and that made me feel as if it wasn't open (weights)*? Please let me know if all are open source or not?
- XenophileJKO 8mo agoI'm assuming the swarm part is all harness. Well I mean a harness and way of thinking that the weights have just been fine tuned to use.
- mccoyb 8mo agoIt's not in the harness today, it's a special RL technique they discuss in https://www.kimi.com/blog/kimi-k2-5.html https://www.kimi.com/blog/kimi-k2-5.html (see "2. Agent Swarm") I looked through the harness and all I could find is a `Task` tool.
- dheera 8mo ago> or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2.5" on the user interface of such product or service. Why not just say "you shall pay us 1 million dollars"?
- 8mo ago
- lrvick 8mo agoActually open source, or yet another public model, which is the equivalent of a binary? URL is down so cannot tell.
- Tepix 8mo agoIt's open weights, not open source.
- typ 8mo agoThe label 'open source' has become a reputation reaping and marketing vehicle rather than an informative term since the Hugging Face benchmark race started. With the weights only, we cannot actually audit that if a model is a) contaminated by benchmarks, b) built with deliberate biases, or c) trained on copyrighted/privacy data, let alone allowing other vendors to replicate the results. Anyways, people still love free stuff.
- Der_Einzige 8mo agoJust accept that IP laws don't matter and the old "free software" paradigm is dead. Aaron Swartz died so that GenAI may live. RMS and his model of "copyleft" are so Web 1.0 (not even 2.0). No one in GenAI cares AT ALL about the true definition of open source. Good.
- duskdozer 8mo agoGood?
- maximgeorge 8mo ago[dead]
- Reubend 8mo agoI've read several people say that Kimi K2 has a better "emotional intelligence" than other models. I'll be interested to see whether K2.5 continues or even improves on that.
- storystarling 8mo agoyes, though this is highly subjective - it 'feels' like that to me as well (comapred to Gemini 3, GPT 5.2, Opus 4.5).
- mohsen1 8mo agoI'll test it out on mafia-arena.com once it is available on Open Router
- Alifatisk 8mo agoYup, I experience the same. I don't know what they do to achieve this but it gives them this edge, really curious to learn more about what makes it so good at it.
- in-silico 8mo agoA lot of people point to the Muon optimizer that Moonshot (the creators of Kimi) pioneered. Compared to the standard optimizer AdamW, Muon amplifies low-magnitude gradient directions which makes the model learn faster (and maybe gives Kimi its unique qualities). Muon paper: https://arxiv.org/abs/2502.16982 https://arxiv.org/abs/2502.16982
- Alifatisk 8mo agoWow! Thank you
- flexagoon 8mo agoI love the Kimi response style. It's much more concise, without all the unnecessary "great question!"s and other annoying AI stuff
- pplonski86 8mo agoThere are so many models, is there any website with list of all of them and comparison of performance on different tasks?
- coffeeri 8mo agoThere is https://artificialanalysis.ai https://artificialanalysis.ai
- pplonski86 8mo agoThank you! Exactly what I was looking for
- XCSme 8mo agoThere are many lists, but I find all of them outdated or containing wrong information or missing the actual benchmarks I'm looking for. I was thinking, that maybe it's better to make my own benchmarks with the questions/things I'm interested in, and whenever a new model comes out run those tests with that model using open-router.
- Reubend 8mo agoThe post actually has great benchmark tables inside of it. They might be outdated in a few months, but for now, it gives you a great summary. Seems like Gemini wins on image and video perf, Claude is the best at coding, ChatGPT is the best for general knowledge. But ultimately, you need to try them yourself on the tasks you care about and just see. My personal experience is that right now, Gemini Pro performs the best at everything I throw at it. I think it's superior to Claude and all of the OSS models by a small margin, even for things like coding.
- Imustaskforhelp 8mo agoI like Gemini Pro's UI over Claude so much but honestly I might start using Kimi K2.5 if its open source & just +/- Gemini Pro/Chatgpt/Claude because at that point I feel like the results are negligible and we are getting SOTA open source models again.
- striking 8mo agohttps://archive.is/P98JR https://archive.is/P98JR
- zmmmmm 8mo agoCurious what would be the most minimal reasonable hardware one would need to deploy this locally?
- NitpickLawyer 8mo agoI parsed "reasonable" as in having reasonable speed to actually use this as intended (in agentic setups). In that case, it's a minimum of 70-100k for hardware (8x 6000 PRO + all the other pieces to make it work). The model comes with native INT4 quant, so ~600GB for the weights alone. An 8x 96GB setup would give you ~160GB for kv caching. You can of course "run" this on cheaper hardware, but the speeds will not be suitable for actual use (i.e. minutes for a simple prompt, tens of minutes for high context sessions per turn).
- simonw 8mo agoModels of this size can usually be run using MLX on a pair of 512GB Mac Studio M3 Ultras, which are about $10,000 each so $20,000 for the pair.
- PlatoIsADisease 8mo agoYou might want to clarify that this is more of a "Look it technically works" Not a "I actually use this" The difference between waiting 20 minutes to answer the prompt '1+1=' and actually using it for something useful is massive here. I wonder where this idea of running AI on CPU comes from. Was it Apple astroturfing? Was it Apple fanboys? I don't see people wasting time on non-Apple CPUs. (Although, I did do this for a 7B model)
- tucnak 8mo agoMac studio way is not "AI on CPU," as M2/M4 are complex SoC, that includes a GPU with unified memory access.
- PlatoIsADisease 8mo agoIf it worked IRL for anything useful, I'd be more interested in the technical differences. But it was a mere toy for a few tests at my fortune 20 company. Language is full of issues of particulars vs universals, and you could debate if its just an integrated GPU with different marketing. Whatever the case, we couldn't use it in production, and NVIDIAs stock price reflects the reality on the ground.
- rvz 8mo agoThe chefs at Moonshot have cooked once again.
- Jackson__ 8mo agoAs your local vision nut, their claims about "SOTA" vision are absolutely BS in my tests. Sure it's SOTA at standard vision benchmarks. But on tasks that require proper image understanding, see for example BabyVision[0] it appears very much lacking compared to Gemini 3 Pro. [0] https://arxiv.org/html/2601.06521v1 https://arxiv.org/html/2601.06521v1
- nostrebored 8mo agoGemini remains the only usable vision fm :(
- Topfi 8mo agoK2 0905 and K2 Thinking shortly after that have done impressively well in my personal use cases and was severely slept on. Faster, more accurate, less expensive, more flexible in terms of hosting and available months before Gemini 3 Flash, I really struggle to understand why Flash got such positive attention at launch. Interested in the dedicated Agent and Agent Swarm releases, especially in how that could affect third party hosting of the models.
- msp26 8mo agoK2 thinking didn't have vision which was a big drawback for my projects.
- bertili 8mo agoThe "Deepseek moment" is just one year ago today! Coincidence or not, let's just marvel for a second over this amount of magic/technology that's being given away for free... and how liberating and different this is than OpenAI and others that were closed to "protect us all".
- motoboi 8mo agoWhat amazes me is why would someone spend millions to train this model and give it away for free. What is the business here?
- testfrequency 8mo agoCurious to hear what “OpenAI” thinks the answer to this is
- YetAnotherNick 8mo agoHosting the model is cheaper per token, the more batched token you get. So they have big advantage here.
- whizzter 8mo agoChinese state that maybe sees open collaboration as the way to nullify any US lead in the field, concurrently if the next "search-winner" is built upon their model the Chinese worldview that Taiwan belongs to China and Tiamen Square massacre never happened. Also their license says that if you have a big product you need to promote them, remember how Google "gave away" site searche widgets and that was perhaps one of the major ways they gained recognition for being the search leader. OpenAI/NVidia is the Pets.com/Sun of our generation, insane valuations, stupid spend, expensive options, expensive hardware and so on. Sun hardware bought for 50k USD to run websites in 2000 are less capable than perhaps 5 dollar/month VPS's today? "Scaling to AGI/ASI" was always a fools errand, best case OpenAI should've squirreled away money to have a solid engineering department that could focus on algorithmic innovations but considering that Antrophic, Google and Chinese firms have caught up or surpassed them it seems they didn't. Once things blows up, those closed options that had somewhat sane/solid model research that handles things better will be left and a ton of new competitors running modern/cheaper hardware and just using models are building blocks.
- pu_pe 8mo agoI don't get this "agent swarm" concept. You set up a task and they boot up 100 LLMs to try to do it in parallel, and then one "LLM judge" puts it all together? Is there anywhere I can read more about it?
- jonkoops 8mo agoThe datacenters yearn for the chips.
- rvnx 8mo agoYou have a team lead that establishes a list of tasks that are needed to achieve your mission then it creates a list of employees, each of them is specialized for a task, and they work in parallel. Essentially hiring a team of people who get specialized on one problem. Do one thing and do it well.
- XCSme 8mo agoBut in the end, isn't this the same idea with the MoE? Where we have more specialized "jobs", which the model is actually trained for. I think the main difference with agents swarm is the ability to run them in parallel. I don't see how this adds much compared to simply sending multiple API calls in parallel with your desired tasks. I guess the only difference is that you let the AI decide how to split those requests and what each task should be.
- zozbot234 8mo agoNope. MoE is strictly about model parameter sparsity. Agents are about running multiple small-scale tasks in parallel and aggregating the results for further processing - it saves a lot of context length compared to having it all in a single session, and context length has quadratic compute overhead so this matters. You can have both. One positive side effect of this is that if subagent tasks can be dispatched to cheaper and more efficient edge-inference hardware that can be deployed at scale (think nVidia Jetsons or even Apple Macs or AMD APU's) even though it might be highly limited in what can fit on the single node, then complex coding tasks ultimately become a lot cheaper per token than generic chat.
- vinhnx 8mo agoOne thing caught my eyes is that besides K2.5 model, Moonshot AI also launched Kimi Code (https://www.kimi.com/code https://www.kimi.com/code), evolved from Kimi CLI. It is a terminal coding agent, I've been used it last month with Kimi subscription, it is capable agent with stable harness. GitHub: https://github.com/MoonshotAI/kimi-cli https://github.com/MoonshotAI/kimi-cli
- Imanari 8mo agoHow does it fare against CC?
- vuldin 8mo agoAnecdotally, I've cancelled my Claude Code subscription after using Kimi K2.5 and Kimi CLI for the last few days. It's handled everything I've thrown at it. It is slower at the moment, but I expect that will improve.
- forgotpwd16 8mo ago>Kimi Code CLI is not only a coding agent, but also a shell. That's cool. It also has a zsh hook, allowing you to switch to agent mode wherever you're.
- vinhnx 8mo agoIt is, Kimi Code CLI supports Zed' Agent Client Protocol (http://agentclientprotocol.com/ http://agentclientprotocol.com/), so it can acts as an external agent that could run in any ACP-compatible client, eg: Zed, Jetbrain, Toad CLI, Minano Notebook. Also, it supports Agent Skills. Moonshot AI developers are actively update the agent and every active. I really like their CLI.
- esafak 8mo agoDoes it support the swarm feature? Does Opencode?
- 8mo ago
- monkeydust 8mo agoIs this actually good or just optimized heavily for benchmarks? I am hopefully its the former based on the writeup but need to put it through its paces.
- kurtis_reed 8mo agoQuite good in my testing
- Barathkanna 8mo agoA realistic setup for this would be a 16× H100 80GB with NVLink. That comfortably handles the active 32B experts plus KV cache without extreme quantization. Cost-wise we are looking at roughly $500k–$700k upfront or $40–60/hr on-demand, which makes it clear this model is aimed at serious infra teams, not casual single-GPU deployments. I’m curious how API providers will price tokens on top of that hardware reality.
- bertili 8mo agoThe other realistic setup is $20k, for a small company that needs a private AI for coding or other internal agentic use with two Mac Studios connected over thunderbolt 5 RMDA.
- zozbot234 8mo agoThat's great for affordable local use but it'll be slow: even with the proper multi-node inference setup, the thunderbolt link will be a comparative bottleneck.
- embedding-shape 8mo agoI'd love to see the prompt processing speed difference between 16× H100 and 2× Mac Studio.
- zozbot234 8mo agoPrompt processing/prefill can even get some speedup from local NPU use most likely: when you're ultimately limited by thermal/power limit throttling, having more efficient compute available means more headroom.
- Barathkanna 8mo agoI asked GPT for a rough estimate to benchmark prompt prefill on an 8,192 token input. • 16× H100: 8,192 / (20k to 80k tokens/sec) ≈ 0.10 to 0.41s • 2× Mac Studio (M3 Max): 8,192 / (150 to 700 tokens/sec) ≈ 12 to 55s These are order-of-magnitude numbers, but the takeaway is that multi H100 boxes are plausibly ~100× faster than workstation Macs for this class of model, especially for long-context prefill.
- hmate9 8mo agoAbout 600GB needed for weights alone, so on AWS you need an p5.48xlarge (8× H100) which costs $55/hour.
- Alifatisk 8mo agoHave you all noted that the latest releases (Qwen3 max thinking, now Kimi k2.5) from Chinese companies are benching against Claude opus now and not Sonnet? They are truly catching up, almost at the same pace?
- zozbot234 8mo agoThe benching is sus, it's way more important to look at real usage scenarios.
- conception 8mo agohttps://clocks.brianmoore.com https://clocks.brianmoore.com K2 is one of the only models to nail the clock face test as well. It’s a great model.
- DJBunnies 8mo agoCool comparison, but none of them get both the face and the time correct when I look at it.
- conception 8mo agoRefresh. It’s not every time but k2 hits a perfect clock for me about 7/10 or so.
- culi 8mo agoKimi 2 is remarkably consistently the best. I wonder if it's somehow been trained specifically on tasks like these. It seems too consistent to be coincidence Also shocking is how the most common runner up I've seen is DeepSeek
- michaelcampbell 8mo agoIt's better than most, but not 100%. As I see this the clock hands are all correct, but the numbers only go 1-8.
- WarmWash 8mo ago
- simonw 8mo agoPretty cute pelican https://tools.simonwillison.net/svg-render#%3Csvg%20viewBox%3D%220%200%20800%20600%22%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%3E%0A%20%20%3Cdefs%3E%0A%20%20%20%20%3ClinearGradient%20id%3D%22skyGradient%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%220%25%22%20y2%3D%22100%25%22%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20style%3D%22stop-color%3A%2387CEEB%3Bstop-opacity%3A1%22%20%2F%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20style%3D%22stop-color%3A%23E0F6FF%3Bstop-opacity%3A1%22%20%2F%3E%0A%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%3ClinearGradient%20id%3D%22beakGradient%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%22100%25%22%20y2%3D%22100%25%22%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20style%3D%22stop-color%3A%23FFB347%3Bstop-opacity%3A1%22%20%2F%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20style%3D%22stop-color%3A%23FF8C00%3Bstop-opacity%3A1%22%20%2F%3E%0A%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%3Cfilter%20id%3D%22shadow%22%20x%3D%22-20%25%22%20y%3D%22-20%25%22%20width%3D%22140%25%22%20height%3D%22140%25%22%3E%0A%20%20%20%20%20%20%3CfeGaussianBlur%20in%3D%22SourceAlpha%22%20stdDeviation%3D%223%22%2F%3E%0A%20%20%20%20%20%20%3CfeOffset%20dx%3D%222%22%20dy%3D%222%22%20result%3D%22offsetblur%22%2F%3E%0A%20%20%20%20%20%20%3CfeComponentTransfer%3E%0A%20%20%20%20%20%20%20%20%3CfeFuncA%20type%3D%22linear%22%20slope%3D%220.3%22%2F%3E%0A%20%20%20%20%20%20%3C%2FfeComponentTransfer%3E%0A%20%20%20%20%20%20%3CfeMerge%3E%0A%20%20%20%20%20%20%20%20%3CfeMergeNode%2F%3E%0A%20%20%20%20%20%20%20%20%3CfeMergeNode%20in%3D%22SourceGraphic%22%2F%3E%0A%20%20%20%20%20%20%3C%2FfeMerge%3E%0A%20%20%20%20%3C%2Ffilter%3E%0A%20%20%3C%2Fdefs%3E%0A%20%20%0A%20%20%3C!--%20Background%20--%3E%0A%20%20%3Crect%20width%3D%22800%22%20height%3D%22600%22%20fill%3D%22url(%23skyGradient)%22%2F%3E%0A%20%20%0A%20%20%3C!--%20Ground%20--%3E%0A%20%20%3Cpath%20d%3D%22M%200%20500%20Q%20400%20480%20800%20500%20L%20800%20600%20L%200%20600%20Z%22%20fill%3D%22%2390EE90%22%2F%3E%0A%20%20%3Cpath%20d%3D%22M%200%20520%20Q%20400%20500%20800%20520%22%20stroke%3D%22%237CFC00%22%20stroke-width%3D%223%22%20fill%3D%22none%22%20opacity%3D%220.5%22%2F%3E%0A%20%20%0A%20%20%3C!--%20Motion%20lines%20--%3E%0A%20%20%3Cg%20opacity%3D%220.3%22%20stroke%3D%22%23666%22%20stroke-width%3D%222%22%20stroke-linecap%3D%22round%22%3E%0A%20%20%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22450%22%20x2%3D%2250%22%20y2%3D%22450%22%2F%3E%0A%20%20%20%20%3Cline%20x1%3D%22120%22%20y1%3D%22470%22%20x2%3D%2260%22%20y2%3D%22470%22%2F%3E%0A%20%20%20%20%3Cline%20x1%3D%22700%22%20y1%3D%22460%22%20x2%3D%22780%22%20y2%3D%22460%22%2F%3E%0A%20%20%20%20%3Cline%20x1%3D%22720%22%20y1%3D%22480%22%20x2%3D%22790%22%20y2%3D%22480%22%2F%3E%0A%20%20%3C%2Fg%3E%0A%20%20%0A%20%20%3C!--%20Bicycle%20Group%20--%3E%0A%20%20%3Cg%20transform%3D%22translate(400%2C%20480)%22%3E%0A%20%20%20%20%3C!--%20Back%20Wheel%20--%3E%0A%20%20%20%20%3Cg%20transform%3D%22translate(-120%2C%200)%22%3E%0A%20%20%20%20%20%20%3Ccircle%20r%3D%2270%22%20fill%3D%22none%22%20stroke%3D%22%23333%22%20stroke-width%3D%228%22%2F%3E%0A%20%20%20%20%20%20%3Ccircle%20r%3D%2260%22%20fill%3D%22none%22%20stroke%3D%22%23DDD%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%3C!--%20Spokes%20--%3E%0A%20%20%20%20%20%20%3Cg%20stroke%3D%22%23AAA%22%20stroke-width%3D%222%22%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%220%22%20y1%3D%22-60%22%20x2%3D%220%22%20y2%3D%2260%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22-52%22%20y1%3D%22-30%22%20x2%3D%2252%22%20y2%3D%2230%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22-52%22%20y1%3D%2230%22%20x2%3D%2252%22%20y2%3D%22-30%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22-30%22%20y1%3D%22-52%22%20x2%3D%2230%22%20y2%3D%2252%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%2230%22%20y1%3D%22-52%22%20x2%3D%22-30%22%20y2%3D%2252%22%2F%3E%0A%20%20%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%0A%20%20%20%20%3C!--%20Front%20Wheel%20--%3E%0A%20%20%20%20%3Cg%20transform%3D%22translate(120%2C%200)%22%3E%0A%20%20%20%20%20%20%3Ccircle% https://tools.simonwillison.net/svg-render#%3Csvg%20viewBox%...
- throwaw12 8mo agoCongratulations, great work Kimi team. Why is that Claude still at the top in coding, are they heavily focused on training for coding or is it their general training is so good that it performs well in coding? Someone please beat the Opus 4.5 in coding, I want to replace it.
- MattRix 8mo agoOpus 4.5 only came out two months ago, and yes Anthropic spends a lot of effort making it particularly good at coding.
- Balinares 8mo agoI replaced Opus with Gemini Pro and it's just plain a better coder IMO. It'll restructure code to enable support for new requirements where Opus seems to just pile on more indirection layers by default, when it doesn't outright hardcode special cases inside existing functions, or drop the cases it's failing to support from the requirements while smugly informing you you don't need that anyway.
- deleted 8mo ago[deleted]
- pokot0 8mo agoI don't think that kind of difference in benchmarks has any meaning at all. Your agentic coding tool and the task you are working on introduce a lot more "noise" than that small delta. Also consider they are all overfitting on the benchmark itself so there might be that as well (which can go in either directions) I consider the top models practically identical for coding applications (just personal experience with heavy use of both GPT5.2 and Opus 4.5). Excited to see how this model compares in real applications. It's 1/5th of the price of top models!!
- symisc_devel 8mo agoGemini 3 pro is way better than Opus especially for large codebases.
- jdeng 8mo agoGlad to to see open source models are catching up and treat vision as first-class citizen (a.k.a native multimodal agentic model). GLM and Qwen models takes different approach, by having a base model and a vision variant (glm-4.6 vs glm-4.6v). I guess after Kimi K2.5, other vendors are going to the same route? Can't wait to see how this model performs on computer automation use cases like VITA AI Coworker. https://www.vita-ai.net/ https://www.vita-ai.net/
- teiferer 8mo agoCan we please stop calling those models "open source"? Yes the weights are open. So, "open weight" maybe. But the source isn't open, the thing that allows to re-create it. That's what "open source" used to mean. (Together with a license that allows you to use that source for various things.)
- Onavo 8mo agoNo major AI lab will admit to training on proprietary or copyrighted data so what you are asking is an impossibility. You can make a pretty good LLM if you train on Anna's Archive but it will either be released anonymously, or with a research only non commercial license. There aren't enough public domain data to create good LLMs, especially once you get into the newer benchmarks that expect PhD level of domain expertise in various niche verticals. It's also a logical impossibility to create a zero knowledge proof that will allow you to attribute to specific training data without admitting to usage. I can think of a few technical options but none would hold water legally. You can use a Σ-protocol OR-composition to prove that it was trained either on a copyrighted dataset or a non copyrighted dataset without admitting to which one (technically interesting, legally unsound). You can prove that a model trained on copywrited data is statistically indistinguishable from one trained on non-copywrited data (an information theoretic impossibility unless there exist as much public domain data as copywrited data, in similar distributions). You can prove a public domain and copywrited dataset are equivalent if the model performance produced is indistinguishable from each other. All the proofs fail irl, ignoring the legal implications, because there's less public domain information, so given the lemma that more training data == improved model performance, all the above are close to impossible.
- deleted 8mo ago[deleted]
- dev_l1x_be 8mo agoI had these weird situations like some models are refusing to use SSH as a tool. Not sure if it was the coding tool limitation or it is baked into in some of the models.
- erichocean 8mo agoRunning on Apple Silicon: https://x.com/awnihannun/status/2016221496084205965 https://x.com/awnihannun/status/2016221496084205965
- stopachka 8mo agoIs there a startup that takes models like this, and effectively gives you a secure setup, where you have (a) a mobile app that (b) talks to some giant machine that only you have access too. If a 10K computer could run this, it may be worth it to have a "fully on prem" version of ChatGPT running for you.
- 2001zhaozhao 8mo agoThe directionally interesting part is that according to the announcement, K2.5 seems to be trained specifically to create sub-agents and work in an agent swarm usefully. The key part is that you don't need to manually create or prompt sub-agents, K2.5 creates them automatically, so from the looks of things it's similar to Claude Code dynamic sub-agents except the model is trained to scale to many more agents autonomously. I wonder whether Claude is doing the same kind of training and it's coming with the next model, and that's why the agent swarm mode in Claude Code is hidden for now. We might be getting very very good agent orchestrators/swarms very soon.
- culi 8mo agoI posted this elsewhere but thought I'd repost here: * https://lmarena.ai/leaderboard https://lmarena.ai/leaderboard — crowd-sourced head-to-head battles between models using ELO * https://dashboard.safe.ai/ https://dashboard.safe.ai/ — CAIS' incredible dashboard * https://clocks.brianmoore.com/ https://clocks.brianmoore.com/ — a visual comparison of how well models can draw a clock. A new clock is drawn every minute * https://eqbench.com/ https://eqbench.com/ — emotional intelligence benchmarks for LLMs * https://www.ocrarena.ai/battle https://www.ocrarena.ai/battle — OCR battles, ELO * https://mafia-arena.com/ https://mafia-arena.com/ — LLMs playing the social deduction game Mafia * https://openrouter.ai/rankings https://openrouter.ai/rankings — marketshare based on OpenRouter
- enricoros 8mo agoCCP-bench has gotten WAY better on K2.5! https://big-agi.com/static/kimi-k2.5-less-censored.jpg https://big-agi.com/static/kimi-k2.5-less-censored.jpg
- raphaelmolly8 8mo ago[dead]