6 ms·
Open source is really about reproducibility. Most of these model releases are better described as "open weights" because we don't know how exactly they were tra
by jallmann 3y ago
Open source is really about reproducibility. Most of these model releases are better described as "open weights" because we don't know how exactly they were trained.
- andy99 3y agoNo, open source is about software freedom, see debian free software guidelines from which the "open source definition" derives. https://wiki.debian.org/DebianFreeSoftwareGuidelines https://wiki.debian.org/DebianFreeSoftwareGuidelines Between these and FSF you've got pretty much all the accepted pontificating about free / open source software. Reproducibility is not mentioned because it's not really a consideration for software. Model weights aren't software so there's not an automatic correspondence between the freedoms, but the essential one you might think you need the training data for is freedom to modify and inspect the source. Modification is fine tuning which you're free to do if you have the weights. And the model weights + code fully define a system that can be interrogated to give a practitioner relevant info about how the model works (within our understanding) that the training data isn't needed or relevant for. I don't see that any freedom on use or inspection is violated by not having the data. It could be nice to have it of course, but it's more about using it to learn how, not exercising any freedom. Incidentally, the big freedom that's usually violated is freedom of discrimination against field of endeavor. LLAMA et al list uses and industries they restrict from using them and because of that are not "open".
- JumpCrisscross 3y ago> open source is about software freedom, see debian free software guidelines from which the "open source definition" The Open Software Foundation ironically screwed the pooch on this one. Open source commonly means source available, more of less. Free software, as in “'free speech,' not as in 'free beer',” is the cumbersome construction for what open source aspired to mean [1]. [1] https://www.gnu.org/philosophy/free-sw.en.html https://www.gnu.org/philosophy/free-sw.en.html
- andy99 3y agoWhy do you say they screwed up? OSI definition is pretty clear. I think the naming is a challenge because in English "open" gives the impression that the key point is that you can see it, as opposed to anything about freedom. Is that what you mean? I have heard it said that open source is sort of a "commercial friendly" version of free software that de-emphasizes user freedom. I think some groups push for that (like Meta is trying to redefine what open source means wrt AI weights). But the OSI defined freedoms basically match what FSF pushes.
- JumpCrisscross 3y ago> Why do you say they screwed up? OSI definition is pretty clear They didn't screw up, they screwed the pooch on open source != source available. The Open Group's members--from IBM to Huawei [1]--started calling the latter open source, which set a precedent that's stuck. [1] https://en.wikipedia.org/wiki/The_Open_Group#Member_Forums_and_Consortia https://en.wikipedia.org/wiki/The_Open_Group#Member_Forums_a...
- monocasa 3y ago> Model weights aren't software so there's not an automatic correspondence between the freedoms, but the essential one you might think you need the training data for is freedom to modify and inspect the source. Models aren't just the weights, but also the list of operations to perform using those weights. I haven't heard a good definition that allows for neural network models to not be software, but allows any other table lookup heavy signal processing algorithm to be software. > Modification is fine tuning which you're free to do if you have the weights. And the model weights + code fully define a system that can be interrogated to give a practitioner relevant info about how the model works (within our understanding) that the training data isn't needed or relevant for. I don't see that any freedom on use or inspection is violated by not having the data. Mistral wouldn't constrain themselves to fine tuning if they have a big enough change, they would go back to their build pipeline. This argument sounds a lot like 'there's nothing stopping you from patching the binary, so that's basically as good as source'.
- btown 3y agoI often like to think about https://github.com/chrislgarry/Apollo-11 https://github.com/chrislgarry/Apollo-11 as an analogy. It's public domain with available source, in the assembly language in which it was written... so it fills all the definitions of OSS! But the process by which that code arose, the ability to modify any line and understand its impact (heh) on a real execution environment, is dependent on a massive process that required billions of dollars and thousands of the smartest people on the planet. For all intents and purposes, without that environment, it is as reliably modifiable as an executable binary in any other context - or a set of weights, in this one!
- monocasa 3y agoI don't think that's a great example. For instance, I can step through and even modify that code using tooling like AGC emulators like this one http://www.ibiblio.org/apollo/#gsc.tab=0 http://www.ibiblio.org/apollo/#gsc.tab=0 What makes it open source is access to the same level of source access that the original developers worked in. That's what's missing here. Mistral's engineers do not simply open this binary in their editor to do their job.
- samus 3y agoReproducibility becomes an important criterion for models though. For normal programs, it is quite easy to decompile an unoptimized binary. Even decompiling an optimized will lead to source code. To make this harder, an obfuscator has to be used. A model is different because it relies on its weight, which are quite a bit more difficult to inspect. Way harder than even obfuscated source code. It is magnitudes harder to make statements about which information it might divulge upon careful questioning, or evaluate its biases, if the training data is not available.
- andy99 3y agoEven if you have the data you can't do that stuff any better. The makeup of the training set doesn't really define the behavior in any tractable way. If anything I think it's a distraction and even when it is available people probably pay too much attention to what's in the training data vs actual behavior. Edit to say that I see benefits to having the training data, just that I don't think it's needed to exercise enough freedom to qualify as open source in an analogous way to software. Also to add, training on GPUs is not generally reproducible anyway because of execution order.
- dizhn 3y agoWould you point to some completely open models if they exist?