5 ms·
> it is not open source It would be nice here if you give some examples of what you call open source model. Please ;) Because the impression is that these thi
by astromaniak 2y ago
> it is not open source
It would be nice here if you give some examples of what you call open source model. Please ;) Because the impression is that these things do not exist, it's just a dream which does not deserve such a nice term..
- Hizonner 2y agoAs far as I know, none have been released. And it doesn't even really make sense, because, as I said, the models aren't copyrightable to begin with and therefore aren't licensable either. However, plenty of open source software exists. The fact that open source models don't exist doesn't excuse attempts to falsely claim the prestige of the phrase "open source".
- kube-system 2y ago> the models aren't copyrightable to begin with What criteria for copyright protection are they missing?
- astromaniak 2y ago> As far as I know, none have been released. I can tell you a secret. What you call 'open source' models are impossible. Because massive randomness is a part of training process. They are not reproducible. Having everything you cannot even tell if the given model was trained on the given dataset. Copyright is a different thing. And a bad news, what's coming is even worst. Those will be the whole things with self awareness and personal experience. They can be copied, but not reproduced. More over, it's hard or almost impossible to detect if something undeclared was planted in their 'minds'. All together means 'open source' model in strict interpretation is a myth, great idea which happen to be not. Like Turing test. > However, plenty of open source software exists. Attempt to switch topic detected. PS: as for that massive downvote, I even wasn't rude, don't care. This account will be abandoned soon regardless, like all before and after.
- jillesvangurp 2y ago> models aren't copyrightable to begin with You are wrong about that. It's a file with numbers. Which makes it a database or dataset and very much protected by copyright. That's why licenses are needed. For the phone book, things like open street maps, and indeed AI models. > The fact that open source models don't exist The fact that many people (myself included) routinely download and use models distributed under OSI approved licenses (Apache V2, MIT, etc.) makes that statement verifiably wrong. And yes, I do check the license of stuff that I use as I work with companies that care about such matters. > As far as I know ... Now you know better.
- Hizonner 2y ago> Which makes it a database or dataset and very much protected by copyright. Not every collection of numbers is a database, and a database is not the same thing as a dataset. Databases have limited copyright-like protection in some places. Under TRIPS, that extends to only databases that are "creative by virtue of the selection or arrangement of their contents" or something along those lines. In the US they talk specifically about curation. ML models do not meet either requirement by any reasonable interpretation. > The fact that many people (myself included) routinely download and use models distributed under OSI approved licenses (Apache V2, MIT, etc.) makes that statement verifiably wrong. The "source code" of an ML model is most reasonably interpreted as including all of the training data, which are never, ever available. Now you know better. [On edit: By the way, the people creating these works had better hope they're outside copyright, because if not, each one of them is a derivative work of (at least some large and almost impossible to identify subset of) its training data, so they need licenses from all the copyright holders of that training material, which few of them have or can get.]
- kube-system 2y agoIf we stop unnecessarily anthropomorphizing software, I think it is plainly obvious these are derivative works. You take the training material, run it through a piece of software, and it produces an output based on that input. Just because the black box in the middle is big and fancy doesn't mean that somehow the output isn't a result of the input. However, transformativeness is a factor in whether or not there is a fair-use exception for the derivative work. And these models are highly transformative, so this is a strong argument for their fair-use.
- simonw 2y agoI'm personally comfortable calling a model "open source" if the license is compatible with the https://opensource.org/ https://opensource.org/ definition. The Llama models aren't. Some of the Mistral models are (the Apache 2 ones). Microsoft Phi-3 is - it's MIT.
- dagaci 2y agoOpen source must include source material so that another can reproduce that the model. I would expect that to be a minimum.
- simonw 2y agoI agree, but that can't happen with the vast majority of these models because they're trained on unlicensed data so they can't slap an open source license on the training data and distribute it. I've decided to draw my personal line at Open Source Initiative compliance for the license they release the model itself under. I respect the opinion that it's not truly open source unless they release the training data as well, but I've decided not to make that part of my own personal litmus test here. My reasoning is that knowing something is "open source" helps me decide what I legally can or cannot do with it when building my own software. Not having access to the training data downs affect my legal rights, it just affects my ability to recompile myself. And I don't have millions of dollars of GPUs so that isn't so important to me, personally.
- Hizonner 2y ago> that can't happen with the vast majority of these models because they're trained on unlicensed data Tough beans? There's lots of actual software that can't be open source because it embeds stuff with incompatible restrictions, but nobody tries to redefine "open source" because of that. ... and, on a vaguely similar-flavored note, you'd better hope that the models you're using end up found to be noninfringing or fair use or something with respect to those "unlicensed data", because otherwise you're in a world of hurt. It's actually a lot easier to argue that the models aren't copyrightable than it is to argue that they're not derivative of the input. > I've decided to draw my personal line at Open Source Initiative compliance for the license they release the model itself under. You're allowed to draw your personal line about what you'll use anywhere you want, but that doesn't mean that you should try to redefine "open source" or support anybody who does.