8 ms·
Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
- jasonjmcghee 2mo ago(2025) As it's 9 months old and they just had a major model release
- GaggiX 2mo agoI believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA), I think previous large Kimi models had only MLA layers.
- yorwba 2mo agoIt's not the same KDA as used in Kimi Linear, though.
- throwa356262 2mo agoFor K3 read this instead: https://arxiv.org/abs/2607.24653 https://arxiv.org/abs/2607.24653 The main contribution of the K3 paper is Stable LatentMoE. Like some other models it compresses data sent between layers, which puts certain requirements on the router. K3 improves performance by using a more balanced expert selection strategy.
- mcbuilder 2mo agoCompared to the Opus 5 "model card", which read like a standard Anthropic set of alignment principles and safety concerns, this presents a plethora of useful technical details that advances the state of the art.
- throwa356262 2mo agoSame with DeepSeek papers, they are a joy to read.
- senko 2mo agoNot an expert, but looks like they did a lot more work on the RL part (9 expert models, full sandbox access for agentic tasks, etc)?
- verdverm 2mo agomost new effort in training comes in the late phase with RL techniques the pretraining (slurping the internet) only goes so far, the new data being used is from human preferences and agent traces (designed and/or distilled)
- cptcobalt 2mo agoRather under-discussed back then: https://news.ycombinator.com/item?id=45766937 https://news.ycombinator.com/item?id=45766937
- Topology1 2mo agoAnother banger from Zhang et. al
- delichon 2mo agoIf you want to believe that the success of Kimi is about distillation attacks, ignore this.
- Parfait__ 2mo agoI stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?
- fwip 2mo agoAnthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen). Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.
- verdverm 2mo agoGoogle is apparently taking a different stance and offering distillation as a paid product https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/distillation https://docs.cloud.google.com/gemini-enterprise-agent-platfo...
- senko 2mo agoOld but relevant: if you read the recently-released Kimi K3 paper[0], you'll see that it's heavily based on Kimi Linear discussed here, scaling it up and adding a bunch more things (like native vision and RL improvements). [0] https://arxiv.org/abs/2607.24653 https://arxiv.org/abs/2607.24653
- bratao 2mo agoI started creating internal models using it, then the Gated Deltanet 2 came out( https://arxiv.org/abs/2605.22791 https://arxiv.org/abs/2605.22791), and it seems like an evolution of it in expressiveness. And in our tests it is really better than.
- iandanforth 2mo agoIs it just me or does this read like a re-implementation of LSTMs?
- muricula 2mo agoI'm no expert but it seems like a descendent of LSTMs. There's a series of papers which show how to reformulate attention as RNNs which arrives at linear attention. Then they add a decay term to get mamba2. Then they add modified the decay term as like a scale to apply both to the existing state and the new update to get delta net. Then they added a gate matrix on the output to get gated delta net. Then Kimi Linear Attention seems to be gated delta net with a more expressive gate. The Gated DeltaNet paper recaptilulates this evolution decently well. But yeah, it feels like they're starting with the same lego blocks and assembling them in similar shapes to accomplish similar but slightly distinct modules.
- throwa356262 2mo agoMore like RNN. The nvidia version is heavy to compute (common problem with RNN and LSTM, also noted in the paper). Moonshot's Kimi K3 replaces part of it with some function that performs better. I dont understand the details but, that in itself might a pretty big contribution. Anyway, I'm amazed how fast these companies improve each other's ideas and put them in new products.
- pooyamo 2mo agoDoes any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can't just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won't beat quicksort if you put more processing to them, but in the modern LLM scene it seems people are in a race to scaling up, experimenting empirically and hoping the same algorithm/architecture comes to a solution.
- chpatrick 2mo agohttp://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- thomasahle 2mo agoHere's one way it could happen: Let's say there's some circuit that does problem solving of the kind we call intelligence. We dont know what this circuit looks like, but it exists in our brain. Doing regression on outputs from the brain (e.g. internet text) with enough parameters, we can "fit" our model to this circuit. But if you try to fit it with fewer parameters than it needs, you're just going to get some linear approximation.
- lacunary 2mo agowhat is the approximation linear in?
- lern_too_spel 2mo agoIn the output of this nonlinear model /s.
- hendiatris 2mo ago
- oakpond 2mo ago>To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints. This is just awesome.
- trollbridge 2mo ago... holy cow!
- Cort3z 2mo agoI believe this is the repo: https://github.com/MoonshotAI/Kimi-Linear https://github.com/MoonshotAI/Kimi-Linear
- imrozim 2mo agoAny one knows how this holds up on long context retrieval (needle in haystack , ruler) vs same size full attention model? efficiency gains look great but that usually where linear attention hybrids fall apart.
- thesiti92 2mo agoif non standard transformers like this take off are companies like etched forked?
- KHUSHIL_GAMER 2mo agoHAKE