4 ms·
Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, proces
by jjordan 1mo ago
Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.
- trvz 1mo agoWhy? Sure, I’d prefer it, too, but this is just another GNU/Linux vs. macOS situation: most of us would prefer the first, but actually get shit done on the latter.
- zufallsheld 1mo agoWithout open-source, there'd be no macOS.. So good thing, it exists.
- homarp 1mo agowhich is why everyone runs docker on mac, to get shit done.
- didibus 1mo agoAnd that's why companies shouldn't fear opening up, but having both is still a net benefit.
- verdverm 1mo agowe get shit done on the cloud with the former rather than the later I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)
- eikenberry 1mo agoWhy do you make the worse choice and not use what you would prefer to use? You have been able to "get shit done" on Linux for nearly 30 years. Have the courage of your convictions.
- kibae 1mo agoThe training data would need to have a permissive license for this to be possible.
- ux266478 1mo agoYou could sidestep it by running non-permissibly licensed training data that you purchased through an LLM. Legal attitude so far seems to be that this is transformative as long as it's not 1:1. The question on whether or not the end result is copyrightable of course remains controversial and inconsistent, but that question is also fairly irrelevent. You don't get more libre than public domain. That's a fair amount of computational and labor overhead mind you, as you'll need to verify and prune the quality of your mountain of synthetic data, but certainly possible. Though this assumes the legal system is a rational actor playing by the set of rules it claims to. In fact, I highly suspect you could get very unlucky and get an unfavorable ruling against you, because you stepped on a big pile of money's toes in the process of doing this.
- alightsoul 1mo agoIt can also be used to sidestep copyright like this forum, books and most websites even if the data was not purchased but is a website or book. Are LLMs what we need to make all data public domain? This way it could be used for that purpose
- echelon 1mo agoEventually we'll just construct 100% synthetic training data that can reliably reproduce pretrains and fine tunes. The first broadly useful fully open source models will do this. We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore. Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works.
- 1mo ago
- cute_boi 1mo agoMoney is the issue here, no one wants to fund it.
- __MatrixMan__ 1mo agoI'm sure anthropic didn't want to fund the extra "safety" guardrails they put into fable, but they were forced to, else they couldn't release it. Sure there are all kinds of problems with that situation. But it still demonstrates that they can be coerced: play nice or don't play at all.
- verdverm 1mo agoI believe Olmo from AllenAi is this https://allenai.org/olmo https://allenai.org/olmo Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech
- yencabulator 1mo agoNo, they just tell you what data they used, it still includes e.g. Common Crawl. Not Open, just willing to state what they fed into the training.
- culi 1mo agoOther than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.
- azinman2 1mo agoThey don’t release all the code.
- culi 1mo agoDeepSeek has released a ton of low-level AI infrastructure code, libraries, and mathematical models on GitHub. Their tools and agent environments are fully open sourced including their harness. They release complete PyTorch and Hugging-face compatible python files detailing their configuration, tokenizers, and layers of the architecture. The only thing they don't release is their data-filtering pipelines but they detail even that in their public-access papers. DeepSeek is truly as open source as you can possibly legally get. Besides the data itself, it's completely reproducible by anyone else. I don't think americans yet acknowledge just how radically transparent Chinese labs are being (and how much even the west benefits from it).
- thepasch 1mo agoInference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open weights models) are still a complete blackbox, and I can't think of any Chinese model that made its training corpus public.
- culi 1mo agoThis is absolutely not true. DeepSeek is most famous for publishing really in-depth papers on their training process but the other labs have started to do the same as well. If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible. In fact everything I just said I said in my original comment. It's like you didn't read it at all. DeepSeek's GRPO Infrastructure, multi-stage training pipeline, and their "cold start" phase have been massively influential in LLM research.
- theplumber 1mo agoBut that would be impossible due copyrights laws. If the law would apply Anthropic and OpenAI executives would be in jail