4 ms·
Absolutely loving this price war, long live open source models.
by sandle 1mo ago
Absolutely loving this price war, long live open source models.
- petcat 1mo ago> long live open source models There are no open source models, at least not useful ones (yet) [0]. Open weight is not the same as open source. The current "open weight" models are just opaque binary blobs you can run on your own computer instead of through a web API. [0] https://allenai.org/ https://allenai.org/
- sleepgummies 1mo ago[dead]
- Matumio 1mo agoThere is also https://apertus-ai.org/ https://apertus-ai.org/ but yeah, "not useful yet" if you were looking to replace your coding agent. Very useful if you are doing LLM research.
- wasfgwp 1mo agoNemotron Super is sort of open source in the sense that Nvidia provides almost everything you need to replicate it from scratch. Of course it’s performance is not exactly stellar but it could be a good starting point for other research teams.
- BoredomIsFun 1mo agoNemotron is okay. Better than Olmo.
- adastra22 1mo agoThis viewpoint doesn't make any sense to me. The weights + inference code are the "source code" for AI. I literally don't know what else you are demanding for the "open source" label.
- keketi 1mo ago> I literally don't know what else you are demanding for the "open source" label Training data and code.
- kingstnap 1mo agoIf you think of LLMs as programs. The weights and inference code are very much a binary. While the training code and data are the true source. Since if you want to robustly modify the LLM that's actually what you need. But since "compilation" (training) is extremely compute intensive this isn't something accessible to anyone without an entire datacenter. Anyway semantics aside having the binary is still infinitely better than dealing with an api as far as privacy and control go.
- kzrdude 1mo agoI agree. But models in difference to compiled binaries, are useful as just weights and can be further refined and post-trained, at least. I don't know LLM theory well enough to say if there's some secret sauce they can hold back that makes training ineffective. Less effective I'm sure, we don't have access to their smart training schemes, but post-training should always be possible IIUC.
- MeetingsBrowser 1mo agoAt the risk of taking the analogy too far, I would treat refining like modifying a dynamic library. You can technically modify behavior, but only in a very coarse way. post-training is like writing a wrapper around the binary. It is closer to building on top of than truly modifying, in that you can tailor things to your needs slightly but cannot make fundamental changes to the underlying thing.
- prometheus1992 1mo agoI really like molmo 2
- drusepth 1mo agoCan you explain what's missing for the "open source" label that open-weight models like DeepSeek/Quen/GLM/etc don't release? Is it just the supplementary data/code for how they were trained, not just the final product?
- petcat 1mo agoThe open weight model providers don't provide the training data or the build tooling. You cannot reproduce the model yourself, or even really know what the model contains. They don't even provide high-level catalogues/descriptions of the training data. An improvement would be something like "trained on the entire WWW up to Aug 1 2026". Or "trained on a Wikipedia archive + Anna's Archive and everything we were able to scrape from Github". They don't provide any of this stuff. I don't mind open-weight models, but they are not open source. It's like bringing home a dog from the rescue and just hoping that it doesn't have a history of biting kids in the face. You just can't know, because you don't know the full history. You can try to add new training (fine tune) to tell it not to bite kids, but that's it.
- drusepth 1mo agoThank you for the detail/explanation. Hope my comment came off correctly as curiosity and a desire to understand, which this helped with.
- rcr-anti 1mo agoWas about to object, then realized you linked Allen ai and had the useful caveat. For what it's worth the OpenMDW license, which Nemotron and a few others have adopted, does say model weight. That said, I've noticed the training procedures and corpus size of more useful open/available weights models are settling down more than I expected. Wonder if crowd sourcing good training data, even if it's just expensive model coding session transcripts, has potential to level the landscape some.