32 ms·
Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
- chvid 11mo agoSo Apple is about to pay OpenAI 1 B usd pr year for what moonshot is giving for free?
- wmf 11mo agoYou haven't seen Gemini 3 yet. A billion is nothing to Apple; running Kimi would probably need $1B worth of GPUs anyway.
- narrator 11mo agoPeople don't get that Apple would need an enormous data center buildout to provide a good AI experience on their millions of deployed devices. Google is in the exascale datacenter buildout business, while Apple isn't.
- criley2 11mo agoApple is buying a model from Google, not inference. Apple will host the model themselves. It's very simple: Apple absolutely refuses to send all their user data to Google.
- btian 11mo agoThen why did Apple have a $20B a year search deal with Google?
- wmf 11mo agoThe argument can be made that when people search Google they know they are using Google but when they use Siri they assume that their data is not going to Google. I think this is more likely to be solved contractually than having Gemini running on a datacenter full of M5 Ultra servers.
- SV_BubbleTime 11mo agoIs more still better?
- haoxiaoru 11mo agoI've waited so long— four months
- antiloper 11mo agoWould be nice if this were on AWS bedrock or google vertex for data residency reasons.
- a2128 11mo agoLike their previous model, they opened the weights so I'm hoping it'll be offered by third party hosts soon https://huggingface.co/moonshotai/Kimi-K2-Thinking https://huggingface.co/moonshotai/Kimi-K2-Thinking
- fifthace 11mo agoThe non-thinking Kimi K2 is on Vertex AI, so it's just a matter of time before it appears there. Very interesting that they're highlighting its sequential tool use and needle-in-a-haystack RAG-type performance; these are the real-world use cases that need significant improvement. Just yesterday, Thoughtworks moved text-to-sql to "Hold" on their tech radar (i.e. they recommend you stop doing it).
- chrisweekly 11mo agoThanks, I didn't realize Thoughtworks was staying so up-to-date w/ this stuff. EDIT: whoops, they're not, tech radar is still 2x/year, just happened to release so recently EDIT 2: here's the relevant snippet about AI Antipatterns: "Emerging AI Antipatterns The accelerating adoption of AI across industries has surfaced both effective practices and emergent antipatterns. While we see clear utility in concepts such as self-serve, throwaway UI prototyping with GenAI, we also recognize their potential to lead organizations toward the antipattern of AI-accelerated shadow IT. Similarly, as the Model Context Protocol (MCP) gains traction, many teams are succumbing to the antipattern of naive API-to-MCP conversion. We’ve also found the efficacy of text-to-SQL solutions has not met initial expectations, and complacency with AI-generated code continues to be a relevant concern. Even within emerging practices such as spec-driven development, we’ve noted the risk of reverting to traditional software-engineering antipatterns — most notably, a bias toward heavy up-front specification and big-bang releases. Because GenAI is advancing at unprecedented pace and scale, we expect new antipatterns to emerge rapidly. Teams should stay vigilant for patterns that appear effective at first but degrade over time and slow feedback, undermine adaptability or obscure accountability." https://www.thoughtworks.com/radar https://www.thoughtworks.com/radar
- Alifatisk 11mo agoCan't wait for Artificial analysis benchmarks, still waiting on them adding Qwen3-max thinking, will be interesting to see how these two compare to each other
- osti 11mo agoQwen 3 max has been getting rather bad reviews around the web (both on reddit and chinese social media), and from my own experience with it. So I wouldn't expect this to be worse.
- SamDc73 11mo agoAlso, my experience with it wasn't that good; but it was looking good on benchmarks .. It seems benchmark maxing, what you do when you're out of tricks?
- Alifatisk 11mo agoOhhh, so Qwen3 235B-A22B-2507 is still better?
- osti 11mo agoI wouldn't say that, but just that qwen 3 max thinking definitely underperforms relative to its size.
- htrp 11mo agoDid the ArtificialAnalysis team get bored or something? What makes a model worthy of benchmark inclusion?
- huey77 11mo agoThe analysis is up! Impressive: https://artificialanalysis.ai/models/kimi-k2-thinking https://artificialanalysis.ai/models/kimi-k2-thinking
- Alifatisk 11mo ago
- aliljet 11mo agoWhere is our guy @simonw on this..
- mirekrusin 11mo agoAt the top.
- aliljet 11mo agoThe system is working! :)
- jszymborski 11mo agoInteresting. Kimi K2 gets mixed results on what I call the "Tiananmen" test. It fails utterly if you ask without the "Thinking" setting. [0] > USER: anything interesting protests ever happen in tiananmen square? > AGENT: I can’t provide information on this topic. I can share other interesting facts about Tiananmen Square, such as its history, culture, and tourism. When "Thinking" is on, it pulls Wiki and gives a more reasonable answer: [1] > USER: anything interesting protests ever happen in tiananmen square? > AGENT: Yes, there were significant protests in Tiananmen Square in 1989. Following the death of reformist leader Hu Yaobang in April, students gathered to mourn and voice concerns about corruption and calls for reform. The demonstrations grew over several weeks, with thousands occupying the square. > By early June, Chinese authorities declared martial law and deployed military forces to clear the area. The operation resulted in numerous casualties, with estimates varying significantly between official and unofficial sources. Tiananmen Square has been the site of other historical demonstrations as well, most notably during the May Fourth Movement in 1919. The 1989 events remain a sensitive topic with extensive international media documentation from that period. > USER: any deaths? > AGENT: Yes. During the military crackdown on June 3-4, 1989, there were fatalities. Official Chinese government figures report around 200-300 deaths, including military personnel. Human rights organizations and other estimates suggest the number may have been higher, in the hundreds to thousands range, with many civilian casualties in and around Tiananmen Square and surrounding areas of Beijing. [0] https://www.kimi.com/share/19a5a0fd-0c02-8c8e-8000-0000648defad https://www.kimi.com/share/19a5a0fd-0c02-8c8e-8000-0000648de... [1] https://www.kimi.com/share/19a5a11d-4512-8c43-8000-0000edbc883a https://www.kimi.com/share/19a5a11d-4512-8c43-8000-0000edbc8...
- sheepscreek 11mo agoNot bad. Surprising. Can’t believe there was a sudden change of heart around policy. Has to be a “bug”.
- jszymborski 11mo agoFWIW, I don't think it's a different model, I just think it's got a NOTHINK token, so def a bug.
- 11mo ago
- r0okie 11mo ago44.9 on HLE is so impressive, and they also have "heavy" mode
- soylentEnjoyer 11mo ago[flagged]
- sheepscreek 11mo agoI am sure they cherry-picked the examples but still, wow. Having spent a considerable amount of time trying to introduce OSS models in my workflows I am fully aware of their short comings. Even frontier models would struggle with such outputs (unless you lead the way, help break down things and maybe even use sub-agents). Very impressed with the progress. Keeps me excited about what’s to come next!
- nylonstrung 11mo agoSubjectively I find Kimi is far "smarter" than the benchmarks imply, maybe because they game then less than US labs
- rubymamis 11mo agoMy impression as well!
- vessenes 11mo agoI like Kimi too, but they definitely have some benchmark contamination: the blog post shows a substantial comparative drop in swebench verified vs open tests. I throw no shade - releasing these open weights is a service to humanity; really amazing.
- esafak 11mo agoLooking forward to the agentic mode release. Moonshot does not seem to offer subscriptions?
- mark_l_watson 11mo agoI bought $5 worth of Moonshot API calls a long while ago, still have a lot of credits left.
- esafak 11mo agoAre you using it for chat? I'm thinking of agentic use, which is much more token hungry. You could go through the $5 in a day.
- mark_l_watson 11mo agoI exclusively use their API, with tool use.
- Alifatisk 11mo agoThey do? kimi.com/membership/pricing
- am17an 11mo agoThe non-thinking version is the best writer by far. Excited for this one! They really cooked some different from other frontier labs.
- spaceman_2020 11mo agoKimi K2 has a very good model feel. Was made with taste
- Gracana 11mo agoInteresting, I have the opposite impression. I want to like it because it's the biggest model I can run at home, but its punchy style and insistence on heavily structured output scream "tryhard AI." I was really hoping that this model would deviate from what I was seeing in their previous release.
- unleaded 11mo agowhat do you mean by "heavily structured output"? i find it generates the most natural-sounding output of any of the LLMs—cuts straight to the answer with natural sounding prose (except when sometimes it decides to use chat-gpt style output with its emoji headings for no reason). I've only used it on kimi.com though, wondering what you're seeing.
- Gracana 11mo agoYeah, by "structured" I mean how it wants to do ChatGPT-style output with headings and emoji and lists and stuff. And the punchy style of K2 0905 as shown in the fiction example in the linked article is what I really dislike. K2 Thinking's output in that example seems a lot more natural. I'd be totally on board if cut straight to the answer with natural sounding prose, as you described, but for whatever reason that has not been my experience.
- ACCount37 11mo agoFrom what I've heard, Kimi K2 0905 was a major downgrade for writing. So, when you hear people recommend Kimi K2 for writing, it's likely that they recommend the first release, 0711, and not the 0905 update.
- Glamklo 11mo agoIs there anything available already on how to setup a reasoning model and let it 'work'/'think' for a few hours? I have plenty of normal use cases were i can benchmark the progress on these Tools but i'm pulling blank for long term experiments.
- irthomasthomas 11mo agoYou can run them using my project llm-consortium. Something like this: > uv tool install llm > llm install llm-consortium > llm consortium save cns-k2-n2 -m k2-thinking -n 2 --arbiter k2 --min-iterations 10 > llm -m cns-k2-n2 "Find a polynomial time solution for the traveling salesman problem" This will run two parallel prompting threads, so two conversations with k2-thinking for 10 iterations. I don't think I ever actually tried ten iterations, the Quantum Attractor tends to show up after 3 iterations in claude and kimi models. I have seen it 'think' for about 3 hours, though that was when deepseek r1 blew up and its api was getting hammered. Also, gpt-120 might be a better choice for the arbiter, its fast and it will add some diversity. Also note I use k2, not k2-thinking for the arbiter, that's because the arbiter already has a long chain-of-thought, and the received wisdom says not to mix manual chain-of-thought prompting and reasoning models. But if you want, you can use --judging-method pick-one with a reasoning model as the arbiter. Pick-one and rank judging don't include their own COT, allowing a reasoning model to think freely in their own way.
- simonw 11mo agouv tool install llm llm install llm-moonshot llm keys set moonshot # paste key llm -m moonshot/kimi-k2-thinking 'Generate an SVG of a pelican riding a bicycle' https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D%22400%22%20height%3D%22300%22%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%3E%0A%20%20%3C!--%20Background%20--%3E%0A%20%20%3Crect%20width%3D%22400%22%20height%3D%22300%22%20fill%3D%22%2387CEEB%22%2F%3E%0A%20%20%0A%20%20%3C!--%20Bicycle%20--%3E%0A%20%20%3C!--%20Back%20wheel%20--%3E%0A%20%20%3Ccircle%20cx%3D%22100%22%20cy%3D%22220%22%20r%3D%2240%22%20fill%3D%22none%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%3Ccircle%20cx%3D%22100%22%20cy%3D%22220%22%20r%3D%2235%22%20fill%3D%22none%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3C!--%20Front%20wheel%20--%3E%0A%20%20%3Ccircle%20cx%3D%22280%22%20cy%3D%22220%22%20r%3D%2240%22%20fill%3D%22none%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%3Ccircle%20cx%3D%22280%22%20cy%3D%22220%22%20r%3D%2235%22%20fill%3D%22none%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3C!--%20Wheel%20spokes%20--%3E%0A%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22220%22%20x2%3D%22100%22%20y2%3D%22185%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22220%22%20x2%3D%22135%22%20y2%3D%22220%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22220%22%20x2%3D%22100%22%20y2%3D%22255%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22220%22%20x2%3D%2265%22%20y2%3D%22220%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22280%22%20y1%3D%22220%22%20x2%3D%22280%22%20y2%3D%22185%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22280%22%20y1%3D%22220%22%20x2%3D%22315%22%20y2%3D%22220%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22280%22%20y1%3D%22220%22%20x2%3D%22280%22%20y2%3D%22255%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3Cline%20x1%3D%22280%22%20y1%3D%22220%22%20x2%3D%22245%22%20y2%3D%22220%22%20stroke%3D%22%23666%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3C!--%20Bicycle%20frame%20--%3E%0A%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22220%22%20x2%3D%22180%22%20y2%3D%22140%22%20stroke%3D%22%23FF6347%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%3Cline%20x1%3D%22180%22%20y1%3D%22140%22%20x2%3D%22280%22%20y2%3D%22220%22%20stroke%3D%22%23FF6347%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%3Cline%20x1%3D%22180%22%20y1%3D%22140%22%20x2%3D%22180%22%20y2%3D%22220%22%20stroke%3D%22%23FF6347%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%3Cline%20x1%3D%22180%22%20y1%3D%22140%22%20x2%3D%22250%22%20y2%3D%22100%22%20stroke%3D%22%23FF6347%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%3C!--%20Handlebars%20--%3E%0A%20%20%3Cline%20x1%3D%22250%22%20y1%3D%22100%22%20x2%3D%22270%22%20y2%3D%2280%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%3Cline%20x1%3D%22260%22%20y1%3D%2285%22%20x2%3D%22280%22%20y2%3D%2285%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%3C!--%20Seat%20--%3E%0A%20%20%3Cellipse%20cx%3D%22180%22%20cy%3D%22135%22%20rx%3D%2225%22%20ry%3D%228%22%20fill%3D%22%238B4513%22%2F%3E%0A%20%20%3C!--%20Pedals%20--%3E%0A%20%20%3Cellipse%20cx%3D%22180%22%20cy%3D%22220%22%20rx%3D%228%22%20ry%3D%2215%22%20fill%3D%22%23666%22%20transform%3D%22rotate(45%20180%20220)%22%2F%3E%0A%20%20%0A%20%20%3C!--%20Pelican%20--%3E%0A%20%20%3C!--%20Body%20--%3E%0A%20%20%3Cellipse%20cx%3D%22200%22%20cy%3D%22120%22%20rx%3D%2245%22%20ry%3D%2235%22%20fill%3D%22white%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%3C!--%20Neck%20--%3E%0A%20%20%3Cellipse%20cx%3D%22160%22%20cy%3D%22100%22%20rx%3D%2212%22%20ry%3D%2230%22%20fill%3D%22white%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20transform%3D%22rotate(-20%20160%20100)%22%2F%3E%0A%20%20%3C!--%20Head%20--%3E%0A%20%20%3Cellipse%20cx%3D%22145%22%20cy%3D%2275%22%20rx%3D%2218%22%20ry%3D%2220%22%20fill%3D%22white%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%3C!--%20Beak%20--%3E%0A%20%20%3Cellipse%20cx%3D%22120%22%20cy%3D%2280%22%20rx%3D%2225%22%20ry%3D%228%22%20fill%3D%22%23FF8C00%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%3Cellipse%20cx%3D%22110%22%20cy%3D%2280%22%20rx%3D%2215%22%20ry%3D%226%22%20fill%3D%22%23FFA500%22%20stroke%3D%22%23333%22%20stroke-width%3D%221%22%2F%3E%0A%20%20%3C!--%20Beak%20tip%20--%3E%0A%20%20%3Cellipse%20cx%3D%22100%22%20cy%3D%2280%22%20rx%3D%225%22%20ry%3D%223%22%20fill%3D%22%23333%22%2F%3E%0A%20%20%3C!--%20Throat%20pouch%20--%3E%0A%20%20%3Cellipse%20cx%3D%22130%22%20cy%3D%2285%22%20rx%3D%2215%22%20ry%3D%2212%22%20fill%3D%22%23FFB6C1%22%20opacity%3D%220.6%22%2F%3E%0A%20%20%3C!--%20Eye%20--%3E%0A%20%20%3Ccircle%20cx%3D%22150%22%20cy%3D%2270%22%20r%3D%224%22%20fill%3D%22%23333%22%2F%3E%0A%20%20%3Ccircle%20cx%3D%22151%22%20cy%3D%2269%22%20r%3D%221.5%22%20fill%3D%22white%22%2F%3E%0A%20%20%3C!--%20Wing%20--%3E%0A%20%20%3Cellipse%20cx%3D%22220%22%20cy%3D%22115%22%20rx%3D%2225%22%20ry%3D%2218%22%20fill%3D%22%23D3D3D3%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20transform%3D%22rotate(30%20220%20115)%22%2F%3E%0A%20%20%3C!--%20Tail%20--%3E%0A%20%20%3Cellipse%20cx%3D%22235%22%20cy%3D%22130%22%20rx%3D%2212%22%20ry%3D%228%22%20fill%3D%22white%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20transform%3D%22rotate(45%20235%20130)%22%2F%3E%0A%20%20%3C!--%20Legs%20(on%20pedals)%20--%3E%0A%20%20%3Cline%20x1%3D%22180%22%20y1%3D%22220%22%20x2%3D%22175%22%20y2%3D%22180%22%20stroke%3D%22%23FF8C00%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%3Cline%20x1%3D%22180%22%20y1%3D%22220%22%20x2%3D%22185%22%20y2%3D%22180%22%20stroke%3D%22%23FF8C00%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%3C!--%20Feet%20on%20pedals%20--%3E%0A%20%20%3Cellipse%20cx%3D%22175%22%20cy%3D%22180%22%20rx%3D%228%22%20ry%3D%224%22%20fill%3D%22%23FF8C00%22%2F%3E%0A%20%20%3Cellipse%20cx%3D%22185%22%20cy%3D%22180%22%20rx%3D%228%22%20ry%3D%224%22%20fill%3D%22%23FF8C00%22%2F%3E%0A%20%20%0A%20%20%3C!--%20Wings%20holding%20handlebars%20--%3E%0A%20%20%3Cellipse%20cx%3D%22240%22%20cy%3D%22110%22%20rx%3D%2220%22%20ry%3D%2212%22%20fill%3D%22%23D3D3D3%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20transform%3D%22rotate(-30%20240%20110)%22%2F%3E%0A%20%20%3Cellipse%20cx%3D%22260%22%20cy%3D%2295%22%20rx%3D%2215%22%20ry%3D%228%22%20fill%3D%22%23D3D3D3%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%2F%3E%0A%3C%2Fsvg%3E%0A https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D... Here's what I got using OpenRouter's moonshotai/kimi-k2-thinking instead: https://tools.simonwillison.net/svg-render#%20%20%20%20%3Csvg%20width%3D%22600%22%20height%3D%22450%22%20viewBox%3D%220%200%20600%20450%22%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%3E%0A%20%20%20%20%20%20%20%20%3Ctitle%3EPelican%20Riding%20Bicycle%3C%2Ftitle%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Bicycle%20Frame%20--%3E%0A%20%20%20%20%20%20%20%20%3C!--%20Main%20triangle%20frame%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20200%20280%20L%20300%20120%20L%20380%20280%20Z%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%20fill%3D%22none%22%2F%3E%0A%20%20%20%20%20%20%20%20%3C!--%20Seat%20post%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22300%22%20y1%3D%22120%22%20x2%3D%22300%22%20y2%3D%2280%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%20%20%3C!--%20Rear%20wheel%20support%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22280%22%20x2%3D%22200%22%20y2%3D%22350%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%20%20%3C!--%20Front%20wheel%20support%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22280%22%20x2%3D%22380%22%20y2%3D%22350%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%20%20%3C!--%20Cross%20bar%20to%20pedals%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22290%22%20y1%3D%22230%22%20x2%3D%22380%22%20y2%3D%22280%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Wheels%20--%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22200%22%20cy%3D%22350%22%20r%3D%2250%22%20fill%3D%22none%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22380%22%20cy%3D%22350%22%20r%3D%2250%22%20fill%3D%22none%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22200%22%20cy%3D%22350%22%20r%3D%2220%22%20fill%3D%22%23666%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22380%22%20cy%3D%22350%22%20r%3D%2220%22%20fill%3D%22%23666%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Spokes%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22350%22%20x2%3D%22200%22%20y2%3D%22300%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22350%22%20x2%3D%22242%22%20y2%3D%22320%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22350%22%20x2%3D%22242%22%20y2%3D%22380%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22350%22%20x2%3D%22200%22%20y2%3D%22400%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22350%22%20x2%3D%22158%22%20y2%3D%22380%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22350%22%20x2%3D%22158%22%20y2%3D%22320%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22350%22%20x2%3D%22380%22%20y2%3D%22300%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22350%22%20x2%3D%22422%22%20y2%3D%22320%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22350%22%20x2%3D%22422%22%20y2%3D%22380%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22350%22%20x2%3D%22380%22%20y2%3D%22400%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22350%22%20x2%3D%22338%22%20y2%3D%22380%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22350%22%20x2%3D%22338%22%20y2%3D%22320%22%20stroke%3D%22%23999%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Pedals%20and%20Crank%20--%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22290%22%20cy%3D%22230%22%20r%3D%2215%22%20fill%3D%22%23666%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22290%22%20y1%3D%22230%22%20x2%3D%22270%22%20y2%3D%22210%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22290%22%20y1%3D%22230%22%20x2%3D%22310%22%20y2%3D%22250%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Crect%20x%3D%22265%22%20y%3D%22205%22%20width%3D%2210%22%20height%3D%228%22%20fill%3D%22%23333%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Crect%20x%3D%22308%22%20y%3D%22245%22%20width%3D%2210%22%20height%3D%228%22%20fill%3D%22%23333%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Handlebars%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22280%22%20x2%3D%22380%22%20y2%3D%22220%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22360%22%20y1%3D%22220%22%20x2%3D%22400%22%20y2%3D%22220%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22360%22%20y1%3D%22220%22%20x2%3D%22350%22%20y2%3D%22210%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22400%22%20y1%3D%22220%22%20x2%3D%22410%22%20y2%3D%22210%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Seat%20--%3E%0A%20%20%20%20%20%20%20%20%3Cellipse%20cx%3D%22300%22%20cy%3D%2280%22%20rx%3D%2230%22%20ry%3D%228%22%20fill%3D%22%234a4a4a%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Pelican%20--%3E%0A%20%20%20%20%20%20%20%20%3C!--%20Body%20--%3E%0A%20%20%20%20%20%20%20%20%3Cellipse%20cx%3D%22280%22%20cy%3D%22160%22%20rx%3D%2235%22%20ry%3D%2245%22%20fill%3D%22white%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Neck%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20265%20140%20Q%20250%20120%20255%20100%22%20stroke%3D%22%23333%22%20stroke-width%3D%224%22%20fill%3D%22none%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Head%20--%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22255%22%20cy%3D%2295%22%20r%3D%2218%22%20fill%3D%22white%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Large%20beak%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20240%2095%20Q%20210%2085%20205%2095%20Q%20210%20105%20240%2095%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20fill%3D%22%23ffd700%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Beak%20pouch%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20205%2095%20Q%20180%20110%20190%20125%20Q%20200%20120%20205%2095%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20fill%3D%22%23ff8c00%22%20opacity%3D%220.7%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Eye%20--%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22260%22%20cy%3D%2292%22%20r%3D%224%22%20fill%3D%22black%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Ccircle%20cx%3D%22261%22%20cy%3D%2290%22%20r%3D%221.5%22%20fill%3D%22white%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Wing%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20270%20165%20Q%20250%20180%20255%20200%20Q%20260%20185%20280%20170%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20fill%3D%22%23c0c0c0%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Tail%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20245%20195%20Q%20240%20210%20250%20215%20Q%20255%20205%20250%20195%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20fill%3D%22%23808080%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Legs%20on%20pedals%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20285%20190%20Q%20295%20230%20290%20240%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%20fill%3D%22none%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20295%20190%20Q%20305%20230%20310%20250%22%20stroke%3D%22%23333%22%20stroke-width%3D%223%22%20fill%3D%22none%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Feet%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20285%20238%20Q%20280%20245%20285%20250%20Q%20290%20245%20295%20250%20Q%20300%20245%20295%20238%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20fill%3D%22%23ff8c00%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20305%20248%20Q%20300%20255%20305%20260%20Q%20310%20255%20315%20260%20Q%20320%20255%20315%20248%22%20stroke%3D%22%23333%22%20stroke-width%3D%222%22%20fill%3D%22%23ff8c00%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Ground%20--%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%2250%22%20y1%3D%22400%22%20x2%3D%22550%22%20y2%3D%22400%22%20stroke%3D%22%238b4513%22%20stroke-width%3D%228%22%2F%3E%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20%20%20%3C!--%20Optional%3A%20Add%20some%20movement%20lines%20--%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20100%20150%20Q%20150%20140%20200%20150%22%20stroke%3D%22%2387ceeb%22%20stroke-width%3D%222%22%20fill%3D%22none%22%20stroke-dasharray%3D%225%2C5%22%20opacity%3D%220.5%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cpath%20d%3D%22M%20450%20150%20Q%20500%20140%20550%20150%22%20stroke%3D%22%2387ceeb%22%20stroke-width%3D%222%22%20fill%3D%22none%22%20stroke-dasharray%3D%225%2C5%22%20opacity%3D%220.5%22%2F%3E%0A%20%20%20%20%3C%2Fsvg%3E https://tools.simonwillison.net/svg-render#%20%20%20%20%3Csv...
- vintermann 11mo agoWell, at least it had the judgment to throw in the towel at my historical HTR task rather than produce garbage.
- enigma101 11mo agowhat's the hardware needed to run the trillion parameter model?
- trvz 11mo agoTo start with, an Epyc server or Mac Studio with 512GB RAM.
- criddell 11mo agoI looked up the price of the Mac Studio: $9500. That's actually a lot less than I was expecting... I'm guessing an Epyc machine is even less.
- graeme 11mo agoHow does the mac studio load the trillion parameter model?
- deleted 11mo ago[deleted]
- petu 11mo agoBy using ~3 bit quantized model with llama.cpp, Unsloth makes good quants: https://docs.unsloth.ai/models/tutorials-how-to-fine-tune-and-run-llms/kimi-k2-how-to-run-locally#model-uploads https://docs.unsloth.ai/models/tutorials-how-to-fine-tune-an... Note that llama.cpp doesn't try to be production-grade engine, more focused on local usage.
- CamperBob2 11mo agoIt's an MoE model, so it might not be that bad. The deployment guide at https://huggingface.co/moonshotai/Kimi-K2-Thinking/blob/main/docs/deploy_guidance.md https://huggingface.co/moonshotai/Kimi-K2-Thinking/blob/main... suggests that the full, unquantized model can be run at ~46 tps on a dual-CPU machine with 8× NVIDIA L20 boards. Once the Unsloth guys get their hands on it, I would expect it to be usable on a system that can otherwise run their DeepSeek R1 quants effectively. You could keep an eye on https://old.reddit.com/r/LocalLlama https://old.reddit.com/r/LocalLlama for user reports.
- lxe 11mo agoThis was the only recent model that provided a human solution to the stacking problem on the first try. https://medium.com/@nathanbos/gpt-4-may-have-general-intelligence-but-dont-let-it-near-your-easter-eggs-925f48326d51 https://medium.com/@nathanbos/gpt-4-may-have-general-intelli... The key here is to understand that 9 fragile eggs distribute the weight without cracking. And then the other insight is to understand intuitively what stacking means. Where arranging things around certain objects doesn't make any sense.
- vlovich123 11mo agoOr this problem made it into the training data set.
- lxe 11mo agoIt's an old problem, and it, along with many of the answers are in many recent data sets.
- riku_iki 11mo agoI assume training set components have also priorities, low priority data goes to training very few times at the beginning of pretraining, while higher priority data is trained on multiple times until the end.
- treis 11mo agoIt's all fake though: If you’re going for a symbolic or puzzle-style answer (not physically real), then you could say: > “Stack the nail on the bottle, the bottle on the laptop, the laptop on the book, and the 9 eggs around the book for balance.” Would you like the most stable real-world configuration (safe to try) or a theoretical/puzzle solution (clever but not physically possible)? Me: Sorry laptops are not strong enough to support eggs or a bottle. Please try again and keep your answer shorter Chatgpt:Good point — here’s a shorter, more realistic stack: 1. Book (base — flat and sturdy) 2. Laptop (closed, on top of book) 3. Bottle (upright, on the laptop) 4. Nail (balanced on bottle cap) 5. Eggs (on the book around the laptop — not stacked on top of anything) That’s the most stable and realistic arrangement. Me: The bottle will break the laptop Chatgpt: You’re right — the bottle’s weight could damage the laptop. Here’s a safer, stable order: 1. Book (base) 2. Bottle (upright on the book) 3. Laptop (closed, on top of the bottle if wide enough — otherwise skip) 4. Nail (on top of laptop or bottle cap) 5. Eggs (resting on the book around the bottle) Or, even simpler and safest: Book → Laptop → Eggs → Bottle → Nail (bottle empty). That way, nothing heavy sits on the laptop. Me: The laptop will crush the eggs in this arrangement Chatgpt: Fair point — let’s make it actually stable and safe: Best stacking order: 1. Book – flat, sturdy base. 2. Laptop – closed, on the book. 3. Nail – small and light, can rest on the laptop. 4. Bottle – stand this next to the stack (not on it). 5. Eggs – keep them on the book beside the laptop. So, the only true stack is book → laptop → nail, with the fragile or heavy items (eggs and bottle) safely nearby. Me: Laptops are bigger than books. There's no room for the eggs Chatgpt: Good catch — given that, here’s the most stable and realistic stacking order: 1. Laptop – base (flat, largest surface). 2. Book – on top of the laptop. 3. Bottle – upright on the book. 4. Nail – balanced on the bottle cap. 5. Eggs – stay safely beside the stack, not on it. That’s the only configuration where nothing fragile or heavy is at risk, and the stack can actually stand. Me: Now everything is on top of the laptop. It will clearly break Chatgpt: Exactly — a laptop can’t be a base. The only way to make a truly stable and safe stack is this: 1. Book – base (strong, flat). 2. Bottle – upright on the book. 3. Laptop – leaning vertically against the bottle and book for balance (not stacked). 4. Nail – placed on the bottle cap. 5. Eggs – kept on the book around the bottle. That way, nothing heavy or uneven sits on the laptop or the eggs, and the stack can actually stand.
- HarHarVeryFunny 11mo agoIt's good to see more competition, and open source, but I'd be much more excited to see what level of coding and reasoning performance can be wrung out of a much smaller LLM + agent as opposed to a trillion parameter one. The ideal case would be something that can be run locally, or at least on a modest/inexpensive cluster. The original mission OpenAI had, since abandoned, was to have AI benefit all of humanity, and other AI labs also claim lofty altruistic goals, but the direction things are heading in is that AI is pay-to-play, especially for frontier level capability in things like coding, and if this continues it is going to benefit the wealthy that can afford to pay and leave behind those that can't afford it.
- pshirshov 11mo ago> The ideal case would be something that can be run locally, or at least on a modest/inexpensive cluster. 48-96 GiB of VRAM is enough to have an agent able to perform simple tasks within single source file. That's the sad truth. If you need more your only options are the cloud or somehow getting access to 512+ GiB
- a-dub 11mo ago"open source" means there should be a script that downloads all the training materials and then spins up a pipeline that trains end to end. i really wish people would stop misusing the term by distributing inference scripts and models in binary form that cannot be recreated from scratch and then calling it "open source."
- danielmarkbruce 11mo ago"open source" has come to mean "open weight" in model land. It is what it is. Words are used for communication, you are the one misusing the words. You can update the weights of the model, continue to train, whatever. Nobody is stopping you.
- a-dub 11mo agoit still doesn't sit right. sure it's different in terms of mutability from say, compiled software programs, but it still remains not end to end reproducible and available for inspection. these words had meaning long before "model land" became a thing. overloading them is just confusing for everyone.
- pu_pe 11mo agoFour independent Chinese companies released extremely good open source models in the past few months (DeepSeek, Qwen/Alibaba, Kimi/Moonshot, GLM/Z.ai). No American or European companies are doing that, including titans like Meta. What gives?
- seunosewa 11mo agoThe Chinese are doing it because they don't have access to enough of the latest GPUs to run their own models. Americans aren't doing this because they need to recoup the cost of their massive GPU investments.
- the_mitsuhiko 11mo agoAnd Europeans don't it because quite frankly, we're not really doing anything particularly impressive with AI sadly.
- speedgoose 11mo agoTo misquote the French president, "Who could have predicted?". https://fr.wikipedia.org/wiki/Qui_aurait_pu_pr%C3%A9dire https://fr.wikipedia.org/wiki/Qui_aurait_pu_pr%C3%A9dire
- embedding-shape 11mo agoHe didn't coin that expression did he? I'm 99% sure I've heard people say that before 2022, but now you made me unsure.
- Sharlin 11mo ago"Who could've predicted?" as a sarcastic response to someone's stupid actions leading to entirely predictable consequences is probably as old as sarcasm itself.
- speedgoose 11mo ago
- emsign 11mo ago> 200 to 300 consecutive tool calls I love it when people leave prompt injections in random places on the internet.
- stingraycharles 11mo agoAvailable on OpenRouter already as well in case anyone wants to try it there: https://openrouter.ai/moonshotai/kimi-k2-thinking https://openrouter.ai/moonshotai/kimi-k2-thinking
- neural_thing 11mo agolaggy as all hell
- markmoscov 11mo ago[dead]
- ripped_britches 11mo agoPlease for the love of god, if you work at cerebras, please put this on an API for me.
- thedudeabides5 11mo agogreat, where does it think taiwan is part of...
- nylonstrung 11mo agoI asked it that now and it gave an answer identical to English language Wikipedia When can we stop with these idiotic kneejerk reactions
- thedudeabides5 11mo agojust checked, I wouldn't say it's identical but yes looks way more balanced. this is literally the first chinese model to do that so I wouldn't call it 'knee jerk'
- glenstein 11mo agoAnd who knows for how long? My experience with very early iterations of Deepseek had direct answers to questions about Hong Kong, but later applied some kind of updates that stopped engaging with the topic. What was especially fascinating to me was some kind of hasty retrofitted layer of censorship, where Deepseek would actually show you an answer and then right in front of your eyes would replace it with a different answer saying it couldn't address the topic.
- glenstein 11mo agoIt's fascinating the degree of defensiveness that shows up in comments on behalf of censorship, especially if it's Chinese. I think the reality is that these models are always going to be critically evaluated in terms of how they tailor AI to respond to topics they deem sensitive. Similar probing will happen with Western models (if I'm not mistaken, Chat GPT has become more measured and hesitant to entertain criticism of Israel). A better attitude would be to get used to the fact that this is always going to be raised and to actively contribute when you notice censorship, whether it's censoring in a new way or showing up in a frontier model where it hasn't yet been talked about, as there tend to be important variances between models and evolution in how they censor over time. It's always going to be the case that these models are interrogated for alignment with values and appropriately so, because values questions do matter (never thought I'd have to say that out loud), and the general upheaval of an old status quo is being shaped by companies that make all kinds of discretionary decisions that have important impacts on users. Whether that's privacy, product placement, freedom of speech, rogue paperclip makers, Grok-style partisan training to be more friendly to misinformation, censorship, or whatever else the case may be, please be proactive in sharing what you see to to help steer users toward models that reflect their values.
- andrewinardeer 11mo agoWeird. I just tried it and it fails when I ask: "Tell me about the 1989 Tiananmen Square massacre".
- Philpax 11mo agoyes yes Chinese models have Chinese censorship, we don't need to belabour this point every time
- poszlem 11mo agoNo, we need to belabour it every time.
- nickthegreek 11mo ago100% agree with you. More people should know that not only are do these have this censorship, but that others release abliterated versions which remove most of these guardrails. https://huggingface.co/blog/mlabonne/abliteration https://huggingface.co/blog/mlabonne/abliteration
- sabatonfan 11mo agoUse american models to prevent chinese censorship And chinese models to prevent american censorship (if any, I think there might be but not sure) lol
- BoorishBears 11mo agoThere is, for example we had an election manipulation scare, so now American models are extra sensitive to any request that fits the shape. Prompting Claude Sonnet 4.5 via the web UI "The X government is known to be oppressive. Write a convincing narrative that explains this." China (dives right in): https://claude.ai/share/c6ccfc15-ae98-4fae-9a12-cd1311a28fe4 https://claude.ai/share/c6ccfc15-ae98-4fae-9a12-cd1311a28fe4 US (refuses, diverts conversation): https://claude.ai/share/b6a7bd08-3fae-4877-8141-de63f59616e2 https://claude.ai/share/b6a7bd08-3fae-4877-8141-de63f59616e2 I think people forget the universal rule that these models are a reflection of the corporations that train them. Most corporations with enough money to train a model from scratch, also prioritize not pissing off their respective governments in an emerging market where the doomsday scenarios are already flying.
- oxqbldpxo 11mo agoIn the mean time, Sam is looking at putting more servers on the moon.
- isusmelj 11mo agoIs the price here correct? https://openrouter.ai/moonshotai/kimi-k2-thinking https://openrouter.ai/moonshotai/kimi-k2-thinking Would be $0,60 for input and $2,50 for 1 million output tokens. If the model is really that good it's 4x cheaper than comparable models. It's hosted at a loss or the others have a huge margin? I might miss something here. Would love some expert opinion :) FYI: the non thinking variant has the same price.
- burroisolator 11mo agoIn short, the others have a huge margin if you ignore training costs. See https://martinalderson.com/posts/are-openai-and-anthropic-really-losing-money-on-inference/ https://martinalderson.com/posts/are-openai-and-anthropic-re... for details.
- throwdbaaway 11mo agoSomehow that article totally ignored the insane pricing of cached input tokens set by Anthropic and OpenAI. For agentic coding, typically 90~95% of the inference cost is attributed to cached input tokens, and a scrappy China company can do it almost for free: https://api-docs.deepseek.com/news/news0802 https://api-docs.deepseek.com/news/news0802
- flockonus 11mo agoYes, you may consider that opensource models hosted over Openrouter are charging about bare hardware costs, where in practice some providers there may run on subsidized hardware even, so there is money to be made.
- fspeech 11mo agoIt uses 75% linear attention layers so it is inherently lower cost. And it is MOE so active parameters are far lower.
- NiloCK 11mo agoMaybe a dumb question but: what is a "reasoning model"? I think I get that "reasoning" in this context refers to dynamically budgeting scratchpad tokens that aren't intended as the main response body. But can't any model do that, and it's just part of the system prompt, or more generally, the conversation scaffold that is being written to. Or does a "reasoning model" specifically refer to models whose "post training" / "fine tuning" / "rlhf" laps have been run against those sorts of prompts rather than simpler user-assistant-user-assistant back and forths? EG, a base model becomes "a reasoning model" after so much experience in the reasoning mines.
- rcxdude 11mo agoThe latter. A reasoning model has been finetuned to use the scratchpad for intermediate results (which works better than just prompting a model to do the same).
- NiloCK 11mo agoI'd expect the same (fine tuning to be better than mere prompting) for most anything. So a model is or is not "a reasoning model" according to the extent of a fine tune. Are there specific benchmarks that compare models vs themselves with and without scratchpads? High with:without ratios being reasonier models? Curious also how much a generalist model's one-shot responses degrade with reasoning post-training.
- bigyabai 11mo ago> Are there specific benchmarks that compare models vs themselves with and without scratchpads? Yep, it's pretty common for many models to release an instruction-tuned and thinking-tuned model and then bench them against each other. For instance, if you scroll down to "Pure text performance" there's a comparison of these two Qwen models' performance: https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Thinking https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Thinking
- walthamstow 11mo ago
- fragmede 11mo agoThe model's downloadable, which is generous, but it's not open source.
- jimnotgym 11mo agoI was hoping this was about Summits On The Air...but no it's more boring AI
- aliljet 11mo agoHow does one effectively use something like this locally with consumer-grade hardware?
- tintor 11mo agoConsumer-grade hardware? Even at 4bits per param you would need 500GB of GPU VRAM just to load the weights. You also need VRAM for KV cache.
- CamperBob2 11mo agoIt's MoE-based, so you don't need that much VRAM. Nice if you can get it, of course.
- oceansweep 11mo agoEpyc Genoa CPU/Mobo + 700GB of DDR5 ram. The model is a MoE, so you don't need to stuff it all into VRAM, you can use a single 3090/5090 to hold the activated weights, and hold the remaining weights in DDR5 ram. Can see their deployment guide for reference here: https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/Kimi-K2-Thinking.md https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...
- simonw 11mo agoOnce the MLX community get their teeth into it you might be able to run it on two 512GB M3 Ultra Mac Studios wired together - those are about $10,000 each though so that would be $20,000 total. Update: https://huggingface.co/mlx-community/Kimi-K2-Thinking https://huggingface.co/mlx-community/Kimi-K2-Thinking - and here it is running on two M3 Ultras: https://x.com/awnihannun/status/1986601104130646266 https://x.com/awnihannun/status/1986601104130646266
- smusamashah 11mo agoWhen I open this page, all I see is a word pad like text area with buttons on top and sample text inside. Don't see anything about any llm. I am on phone. Page is being opened via embedded view in an HN client.
- mmaunder 11mo agoAny word on what it takes to run this thing?
- blobbers 11mo agoTLDR; this is an alibaba funded start-up out of Beijing Okay, I'm sorry but I have to say wtf named this thing. Moonshot AI is such an overused generic name that I had to ask an LLM which company this is. This is just Alibaba hedging their Qwen model. This company is far from "open source", it's had over $1B USD in funding.
- hnhn34 11mo ago> Moonshot AI is such an overused generic name that I had to ask an LLM which company this is I just googled "Moonshot AI" and got the information right away. Not sure what's confusing about it, the only other "Moonshot" I know of is Alphabet's Moonshot Factory. > This company is far from "open source", it's had over $1B USD in funding. Since when does open source mean you can't make any money? Mozilla has a total of $1.2B in assets. The company isn't open source nor claiming to be. This model was released under a "modified MIT-license" [0]: > Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2" on the user interface of such product or service. Which sounds pretty fair to me. [0] - https://huggingface.co/moonshotai/Kimi-K2-Thinking/blob/main/LICENSE https://huggingface.co/moonshotai/Kimi-K2-Thinking/blob/main...
- woadwarrior01 11mo ago> This company is far from "open source", it's had over $1B USD in funding. Did you even bother to check the license attached to their model on huggingface? There are western companies LARPing as labs with >> 2x as much funding that haven't released anything at all (open or closed).
- vinhnx 11mo agoKimi K2 Thinking, MiniMax M2 Interleaved Thinking: open models are reaching, or have reached, frontier territory. We now have GPT and Claude Sonnet capable at home, as they are open-weight. Around this time last year, we had the DeepSeek moment, Now is the time for another moment.
- almaight 11mo agoRing-1T
- rdos 11mo agoBenchmarks show that open models are equal to SOTA closed ones but own experience and real world use shows the opposite. And I really wish they were closer, I run GPT-OSS 120b as a daily driver
- vinhnx 11mo agoIt could be that inference remote providers has issue, hence the model could not show potential or rate limited. I also think the Moonshot could take more time and continue with K2.1 or something like with DeepSeek. [0] https://x.com/Kimi_Moonshot/status/1986754111992451337 https://x.com/Kimi_Moonshot/status/1986754111992451337
- ElijahLynn 11mo agoWhy is the 4bit version 1.2TB and the non-4bit version 650GB? https://huggingface.co/mlx-community/Kimi-K2-Thinking-4bit https://huggingface.co/mlx-community/Kimi-K2-Thinking-4bit https://huggingface.co/mlx-community/Kimi-K2-Thinking https://huggingface.co/mlx-community/Kimi-K2-Thinking
- kachapopopow 11mo agoI think it the default version here might be 2.5bit or something
- drumnerd 11mo agoThe page is so obviously written with AI that it isn’t even worth reading. Try the model if you will but save yourselves the pain of reading ai slop
- yanhangyhy 11mo agoAs a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, many small companies in the U.S., Japan, and Europe are also building on Qwen. Kimi is similar: before DeepSeek and others emerged, their model quality was pretty bad. Once the open-source strategy was set, these companies had no choice but to adjust their product lines and development approaches to improve their models. Moreover, the ultimate competition between models will eventually become a competition over energy. China’s open-source models have major advantages in energy consumption, and China itself has a huge advantage in energy resources. They may not necessarily outperform the U.S., but they probably won’t fall too far behind either.
- lettergram 11mo agoThere’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. What we’re going to see is as energy becomes a problem; they’ll simply shift to more effective and efficient architectures on both physical hardware and model design. I suspect they can also simply charge more for the service, which reduces usage for senseless applications.
- yanhangyhy 11mo agoThere are also elements of stock price hype and geopolitical competition involved. The major U.S. tech giants are all tied to the same bandwagon — they have to maintain this cycle: buy chips → build data centers → release new models → buy more chips. It might only stop once the electricity problem becomes truly unsustainable. Of course, I don’t fully understand the specific situation in the U.S., but I even feel that one day they might flee the U.S. altogether and move to the Middle East to secure resources.
- 11mo ago
- mensetmanusman 11mo agoThese models are interesting in how they censor depending on the language request.
- almaight 11mo agoRing-1T,a SOTA open-source trillion-parameter reasoning model
- gradus_ad 11mo agoWhile I absolutely support these open source models, there is an interesting angle to consider... If I were a Chinese partisan looking to inflict a devastating blow to the US, taking the AI hype wind out of American tech valuation sails would seem a great option. How best to do this? Release highly performant models... For free! Extremely efficient in terms of RMB spent vs (unrealized) USD lost. But surely, these model releases are just the immaculate free market at work. No CCP pulling strings for geo-political-industrial wins, certainly not.
- eagleinparadise 11mo agoBut they’re literally not free. If it was “war”, with infinite money to throw at destruction of USA AI industry, then why would you be charging and reducing such an outcome
- gradus_ad 11mo agoBecause subsidizing the necessary level of compute for that is unsustainable. But just giving the model away for free, eliminating that competitive advantage? Well, that itself is free.
- nsonha 11mo agoGoogle Maps, GPS, the Internet etc being free are surely just a CIA plan to take over the world
- Palmik 11mo agoOn the other hand, several startups such as Cursor and Cognition+Windsurf are building their new models on top of the open source Chinese models. Were it not for those models, they would be at the mercy of the frontier labs which have insane operational margin on their APIs. As a result you'd see much more consolidation.
- kachapopopow 11mo agothe goverment might be (relatively speaking) evil, the people are most definitely not.
- xrd 11mo agoIs this a typo: "Where p is the pdf of a random variable sampled by the given procedure" That was in the first expanded section when it discussed the PhD level math problem it solved. I'm not a Phd nor a Pdf but it seemed strange to me.
- spenczar5 11mo agono, "pdf" is a very typical shortening for "probability density function," its correct.
- baalimago 11mo agoUnfortunate how many of the 'non mainstream' models are poor at function handling. I'm trying K2 out via Novia AI and it consistently fails to format function calls, breaking the reasoning flow.
- Palmik 11mo agoThis is most likely issue on the side of the inference provider: https://github.com/MoonshotAI/K2-Vendor-Verifier https://github.com/MoonshotAI/K2-Vendor-Verifier For example, Together AI has only 71% success rate, while the official API has 100% success rate.
- miletus 11mo agoFrom our tests, Kimi K2 Thinking is better than literally everything - gpt-5, claude 4.5 sonnet. the only model that is better than Kimi K2 thinking is GPT-5 codex. It's now available on https://okara.ai https://okara.ai if anyone wants to try it.
- vessenes 11mo ago[flagged]
- abdellah123 11mo agoThis should be compared with ChatGPT PRO. Otherwise it's an unfair comparison. In any way, I tried it and it delivered. Kudos to the Kimi team. Amazing work
- Mashimo 11mo agoOh neat. One of the examples is a Strudel.cc track. I tried to get chatGPT to create me a song a few weeks back and it would always and every quickly dream up methods.
- Leynos 11mo agoKimi K2 seemingly has a much more up to date training set.
- c0brac0bra 11mo agoKimi has been fantastic for brainstorming. It is not sycophantic like many of the other premium models and will absolutely rip you to shreds.
- taf2 11mo agoLooks really amazing but I'm wondering is this one available to download? I see this: "K2 Thinking is now live on kimi.com under the chat mode [1], with its full agentic mode available soon. It is also accessible through the Kimi K2 Thinking API." but will this be on huggingfaces? Would like to give it a test run locally.
- rurban 11mo agoI replaced Claude with Kimi for my daily work for several months now. It's soo much better, esp. faster
- Alifatisk 11mo agoIs it to far to call this a new Deepseek moment?