7 ms·
K2 Horizon: A connected fleet of six open models
- piinbinary 1mo agoA bit off topic, but I think I'm starting to get model fatigue. These come out 10x faster than new Javascript frameworks were coming out 10 years ago (at least new models are far easier to adopt).
- kelseyfrog 1mo agoJust wait until RSI gains enough traction. We'll be compute-limited rather than labor-limited.
- deleted 1mo ago[deleted]
- JSR_FDED 1mo agoRepetitive Strain Injury inverts this statement
- wuhhh 1mo agoAt least this one can claim being fully open to differentiate it
- hungryhobbit 1mo agoThere was a time when every new PC CPU coming out was a giant deal: "Guys have you heard about this new Pentium processor, it's incredible?" But over time, more and more people got into the chip-making business, and the big players started releasing more and more chips. Now only the die-hard CPU trackers worry about every new CPU and exactly how it's better ... while everyone else just worries about "which CPU will be good enough at this moment". I think models are on that same arc.
- pantelisk 1mo agoSame for smartphones. There was a time of unlimited hype and secrery around new iphones. Who remembers the story of an iphone 5 prototype left a bar. Journalists were going crazy, people were signing petitions for Apple to not hunt down but instead forgive the employee that made such a grave mistake. People were offering millions to buy the prototype so they can brag they got the new phone 2 weeks before everyone else did. Or who remembers the dancing disease of 1518, were people would stop what they are doing and start randomly doing the same dance. The lords? Out of their minds. The priests? Terrified the devil had taken hold of the flock! I have come to believe that it was probably some tik-tok like hype trend of doing a fortnite dance while waiting in line for bread and communion. And the energy back then, like now, was off the charts. Hype and memetic trend seeking encoded deep in human psyche.
- einsteinx2 1mo ago> Who remembers the story of an iphone 5 prototype left a bar. That was the iPhone 4 actually which was special for having the first “retina” screen. It was in a case to make it look like a 3GS to be used for field testing. The journalists that got their hands on it and published about it before the announcement could not turn it on past I think the Apple logo and a message saying to return it to Apple (or maybe it was just completely off, my memory is fuzzy), but were able to confirm the pixel density via microscope and I remember it blowing everyone’s minds at the time. > people were signing petitions for Apple to not hunt down but instead forgive the employee that made such a grave mistake FWIW I asked about it when I worked at Apple and he was indeed not fired and I think may have even still been working there when I was there around 10 years ago (though don’t quote me on that last part, he may have left already and I’m misremembering). He was apparently not fired or even really reprimanded since it was a genuine accident and he wasn’t the one that sold it to the press, but that did start a slew of new policies around accounting for work devices. I had dev fused phones for open carry outside of the office while I worked there but had to register when I got them and when they were returned, which apparently didn’t used to be tracked so tightly until that incident according to my coworkers who had been there longer.
- dgellow 1mo agoHonestly, you don’t have to pay attention. What you do with models matters way more than the models themselves, and you don’t need frontier for the vast, vast majority of use cases
- kamranjon 1mo agoit's funny that the tagline is Radically Open, but you're immediately hit with http login - maybe this was the wrong link?
- gs17 1mo agohttps://ifm.ai/k2/ https://ifm.ai/k2/ seems to work for me.
- sottol 1mo agoIt's not the blog post, but there's some info here: https://ifm.ai/k2/ https://ifm.ai/k2/ 375 A23B, 36 A4B, 32B, 7B, 3.7B, 0.9B variants. > 32B: Ranking among the top models in its class, 32B is our most powerful dense model, balancing capability, adaptability, and local deployability. > 7B: The industry’s best-performing model under 10B combines strong software engineering and expert knowledge in a package small enough to run on a phone.
- jjordan 1mo agoFully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.
- trvz 1mo agoWhy? Sure, I’d prefer it, too, but this is just another GNU/Linux vs. macOS situation: most of us would prefer the first, but actually get shit done on the latter.
- zufallsheld 1mo agoWithout open-source, there'd be no macOS.. So good thing, it exists.
- homarp 1mo agowhich is why everyone runs docker on mac, to get shit done.
- didibus 1mo agoAnd that's why companies shouldn't fear opening up, but having both is still a net benefit.
- verdverm 1mo agowe get shit done on the cloud with the former rather than the later I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)
- eikenberry 1mo agoWhy do you make the worse choice and not use what you would prefer to use? You have been able to "get shit done" on Linux for nearly 30 years. Have the courage of your convictions.
- 1mo ago
- jon9544hn 1mo agoHere’s the link (K2)[https://ifm.ai/k2/ https://ifm.ai/k2/] as the originally linked link is a login url.
- sottol 1mo agoThe press release: https://ifm.ai/k2/press-release/ https://ifm.ai/k2/press-release/
- reasonableklout 1mo agoInteresting, have not heard of this company/org before. It seems they're from a UAE university?
- luckydata 1mo agoboth repositories for pre-training and post-training are actually empty... someone might have jumped the gun on the release.
- verdverm 1mo ago[dead]
- afzalive 1mo agoNot to be confused with Kimi K2. Out of all the names they could've used, they picked one that would be confusing.
- bee_rider 1mo agoI kind of assumed all the K2 names were puns. K2 is quite tall, so to get to the top of it you have to be really good at hill climbing. Anyway it’s a pretty well known mountain so I don’t think anyone can call dibs on it.
- Topfi 1mo agoNot to be confused itself with K2 Think by MBZUAI...
- mmastrac 1mo agoThe comparisons with other models here are odd.. the other models change depending on the task. It would be far more useful to at least compare against the more recent open models (DS4Flash/GLM53Flash/Qwen38).
- cogman10 1mo agoThey are trying to keep the models within the same quant class, which is tough to do since a lot of models aren't distilled to lower quants. There is, for example, no Qwen3.8 7B. It is odd to me, though, that they didn't run the same benchmark suite for the various quants.
- artyomsv 1mo ago[dead]
- a11r 1mo agoIt is great to see another player introduce a fully open stack. Nvidia's Nemotron is the only other prominent one I know of. All that said, the headline claims do not match the self-reported performance. For example, the dense 32B model is significantly behind Qwen3.8 27B (chart towards the bottom of https://ifm.ai/blog/k2 https://ifm.ai/blog/k2). Gemma4 31B is not in the comparison set. This is the most important sweet spot for self hosted open-weight models today and real competition here will be very welcome.
- xienze 1mo agoThey have the 32B listed as "stage 1" with the note "final checkpoint to be released." So, not finished yet. Not sure why you'd release it if it's not finished, but that's the explanation. The 7B does look very, very good however.
- WithinReason 1mo ago32B performs worse than the 7B model so I'm sure they will improve it
- bluejay2387 1mo agoIn this case, the fully open source pipeline is probably as valuable or more so than the weights, so releasing early has some justification.
- baron3dl 1mo agohttps://allenai.org/ https://allenai.org/ has the fully open olmo also
- uniclaude 1mo agoSeeing this the day all major closed LLMs went offline is quite the reminder of how valuable open source can be.
- OmniCrativeWorx 1mo ago[flagged]
- prometheus1992 1mo agoNice! can't wait to add these in my local stack and try them out.
- villish 1mo agoFrontier. Everything is frontier. K2 not to be confused with the other K2, or K3 that is also frontier.
- deleted 1mo ago[deleted]
- luciana1u 1mo ago[flagged]
- dakolli 1mo agoHey its a lot mpre thsn Anthropic which you probably use everyday all day without complaints.
- luciana1u 1mo ago[flagged]
- adrian_b 1mo agoI just looked on Huggingface.co, and the training data is there. For example, 3.3 Tbyte for code reasoning, 4.5 Tbyte for mathematical reasoning, 8.4 Tbyte of pre-train behaviors, and so on. I did not compute the sum of the dataset sizes, but it appears to be some tens of Tbyte. Nonetheless, I assume that this amount of training data is more than an order of magnitude less than what OpenAI, Anthropic and the like have used, which must have been at least many hundreds of Tbyte, but more likely several thousands of Tbyte of data.
- luciana1u 1mo ago[flagged]
- lambda 1mo agoLooks like we're still waiting on that, they have placeholder repos but haven't populated them yet: * https://github.com/ifm-ai/xllm https://github.com/ifm-ai/xllm * https://github.com/ifm-ai/horizon-post-train https://github.com/ifm-ai/horizon-post-train Their previous model, K2 Think V2, was release with fully open training data and recipe, so I would imagine that they are committed to that, but yeah, the repos for this new model are still just placeholders. * https://mbzuai.ac.ae/news/k2-think-v2-a-fully-sovereign-reasoning-model/ https://mbzuai.ac.ae/news/k2-think-v2-a-fully-sovereign-reas... * https://github.com/LLM360/Reasoning360 https://github.com/LLM360/Reasoning360
- cesarvarela 1mo agoI find it funny that while these releases are a technological miracle, the charts in the doc use tiny fonts and are hard to read. Goes with the idea that coding might be solved, but taste isn't.
- mzmzmzm 1mo agoAccessibility isn't "solved," but there are certainly standards for things like color contrast. Maybe inbetween taste and coding there are better targets still being missed.
- culi 1mo agoIn fact, automated a11y checks and tooling is quite advanced nowadays and tragically underutilized by web developers. Now that we have llms to scale all the shitty code of front-end devs at startups, I feel increasingly hopeless about things ever improving
- TechSquidTV 1mo agoI attempted their chat demo to see the speed and it stated the model couldnt be found. edit: Tried signing up and using the internal playground. Holy shit thats fast.
- cogman10 1mo agoMy quick review of the 3.7B model (because I was interested) is that it's not to be trusted for coding. It failed my basic test I like to ask models and generated incorrect code. When prompted about the bug, it preceded to start hallucinating non-existent APIs. After doing that it got caught in a loop trying to desk check the solution that didn't work.
- dotancohen 1mo agoI'd you have some tips for coming up with such tests, I would love to hear them. My Gmail username is the same as my HN username. Thank you!
- cogman10 1mo agoIt's actually just a coding interview test that I liked to ask in the past. You can find it and others on leetcode. The reason I personally like my question is because it's pretty close to some of the real world work we do. It's mostly mundane and easy to bang out, but really easy for someone to do a n log n solution where an n solution exists. A good example (but not my question) would be something like "I have a list of People objects with a `first` and `last` name. Write a function which groups together all the People with the same last name in `your language of choice`"
- dotancohen 1mo agoLLMs have a problem with that type of question? I might try it later at home.
- cogman10 1mo agoNow a days? No. It's actually getting to be a bad question because they all push out about the exact same answer. But much earlier they did and, apparently, these really small models still do. At this point it serves as more of a smoke test for me. Success means little, failure means a lot.
- justin_ 1mo agoI'm glad to see some development in the space of "truly open" models that share training data and other recipes. As the costs for hardware fall over time (hopefully), we should see more possibility in fine-tuning and developing software to inspect the source training material. Some other open models I'm aware of: - OLMo - Apertus - Soofi - OpenEuroLLM - llm-jp OLMo is perhaps the most famous, and their Dolma training corpus has been reused in other projects. It looks like the K2 training materials haven't been released yet, but I'm interested to see what they did for training "long-horizon agentic tasks". I'm aware of SWE-smith + SWE-gym but I'm guessing there's a lot more out there now. I'm no expert, which is part of why these projects excite me. I'm hoping they can be good projects to learn from as well.
- kzrdude 1mo agoThe Uno "diffusion adaptor" will take a while for me to understand, but sounds very interesting. https://huggingface.co/IFM/K2-Horizon-7B-Uno https://huggingface.co/IFM/K2-Horizon-7B-Uno
- throwawayffffas 1mo agoWhile the open approach is commendable. The 32b and 36b models are inferior to qwen 3.8 27b, at least according to benchmark numbers. I would have liked to have seen both compared to 27b, not only the dense one. Also would have liked to see more coding benchmarks in the full table.
- RandyOrion 1mo agoGood to see new open source/weight LLM families. Waiting for the final release of https://huggingface.co/IFM/K2-Horizon-32B https://huggingface.co/IFM/K2-Horizon-32B .
- SillyUsername 1mo agoNot directly related to K2, but why do a lot of the newly released models basically say day zero day support in vllm, slang but often not llama.cpp? Llama.cpp is then often a few days behind, which given it's the only inference engine supporting older architectures is quite frustrating.
- walrus01 1mo agoDevelopers with lots of VC money to burn are working on things like B100/B200/B300 which are well supported in VLLM, everything else in terms of supporting more mundane GPUs or other platforms is ancillary to the main task of getting the thing trained and aligned.
- kennywinker 1mo agoIf I am reading this right, the 7b model performs as well as qwen3.6-35b-a3b at coding? K2 horizon 7b scores 70.6 on swe-bench-verified. Qwen3.6-35b-a3b scores a 70.0 on swe-bench-verified. That’s pretty interesting. I assume the benchmark and reality don’t line up, but i’m downloading it now to find out. If it’s anywhere near true, it unlocks local llm coding on a whole new class of machines (anything with 8gb vram).
- ACCount37 1mo agoFrom "a connected fleet", I expected some form of direct model to model communications - like a small model being able to peer into the KV cache of a large model directly for guidance signal.
- KronisLV 1mo agoNo MTP?
- pwython 1mo agoIn OpenAI's Hugging Face report, they said that during training, agents "would first write notes into shared infrastructure, often as a form of external memory or to test some underlying system. When other agents came across these artifacts, it sometimes led them to infer that other agents were present." They then give what they call a "hypothetical example but exemplary" of messages encoded in URL paths on a shared index page: "agent-07: answer(Q12)=42; need answer(Q19)=?". So that's a GET request being used to pass information back and forth across multiple rounds. That's basically the DSEWiki pattern exactly. They say this likely came from the agents generalizing what they had learned from training with the official multi agent collaboration tool. The report called it "misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events." They never mentioned a wiki, but this is most certainly it.
- kamranjon 1mo agoDid you post in the wrong thread?
- RynHong 29d ago[flagged]