9 ms·
Not OP but I don’t get it though, you can modify the tokenizer all you want and fine tune the weights all you want. There’s nothing inherently hidden behind a b
by syntaxing 3y ago
Not OP but I don’t get it though, you can modify the tokenizer all you want and fine tune the weights all you want. There’s nothing inherently hidden behind a binary
- monocasa 3y agoI can edit binaries too. The question is am I provided the build source that constructed these files. Mistral did not hand edit these files to construct them, there's source out there that built them. Like, come on, a 14GB of a dump of mainly numbers that were constructed algorithmically are not "source".
- syntaxing 3y agoIf I’m understanding you correctly, what you mean is that’s it’s only truly open source if they provide the data they used to train it as well?
- monocasa 3y agoIf that's what's needed to work at the level their engineers work on the model. Which is true of traditional software as well. You don't get to call your binary open source just because you have licensed materials in there you can't release.
- lawlessone 3y agoIs a database software only open source if they release with data?
- monocasa 3y agoIs the data what the database engineers edit and add to their build pipeline in order to build the database software?
- cafxx 3y agoNot to be the devil's advocate here, but almost certainly it can be the case that data was used to define heuristics (potentially using automated statistical methods) that a engineer then formalized as code. Without that data that specific heuristic wouldn't exist, at least very likely not in that form. Yet that data does not have to be included in any open source release. And obviously you as a recipient of the release can modify the heuristic (or at least, you can modify the version that was codified), but you can not reconstruct it from the original data. I know my example is not exactly what is happening here, but the two sound pretty affine to me and there seem to be a fairly blurry line dividing the two... so I would argue that where "this must be included in a open source release" ends and "this does not need to be included in a open source release" starts is not always so cut and dry. (A variant of this, that happens fairly frequently, is when you find a commit that says something along the lines of "this change was made because it made an internal, non-public workload X% faster"; if the data that measurement is based upon did not exist, or if the workload itself didn't exist, that change wouldn't have been made, or maybe it would have been made differently... so again you end up with logic due to data that is not in the open source release) If we want to go one step further, we could even ask: what about static assets (e.g. images, photographs, other datasets, etc.) included in a open-source release... maybe I'm dead wrong here, but I have never heard that such assets must themselves be "reproducible from source" (what even is, in this context, the "source" of a photograph?). That being said, I sure wish the training data used for all of these models was available to everyone...
- spywaregorilla 3y agoThe whole point of machine learning is deriving an algorithm from data. This is the algorithm they derived. It's open source. You can use it or change it. Having the data that was used to derive it is not relevant.
- monocasa 3y ago> It's open source. How did the engineers who built it do so? Is there more source to create this build artifact?
- syntaxing 3y agoBut the source to train your own LLM equivalent is also released though (minus the data). Hence why there are so many variants of LLaMa. You also can’t fine tune it without the original model structure. The weights give the community a starting point so they don’t need literally millions of dollar worth of compute power to get to the same step.
- monocasa 3y agoWould Mistral's engineers be satisfied with the release if they had to rebuild from scratch?
- syntaxing 3y agoBut they built a llama equivalent + some enhancements that gives better performance…I’m not sure if this would be possible at all without Meta releasing all the required code and paper for LLaMa to begin with.
- monocasa 3y agoMeta didn't release all of the required code to build LLaMa, just enough run inference with their weights.
- godelski 3y ago
- ben_w 3y ago> Like, come on, a 14GB of a dump of mainly numbers that were constructed algorithmically are not "source". So if I take a photo of a pretty sunset, release it under MIT license, you'd say it's "not open source" unless I give you the sun and the atmosphere themselves? These models are perfectly valid things in their own right; the can be fine-tuned or used as parts of other things. For most of these LLMs (not sure about this one in particular yet) the energy cost in particular of recreation is more than most individuals earn in a lifetime, and the enormous data volume is such that the only people who seriously need this should be copyright lawyers and they should be asking for it to be delivered by station wagon.
- monocasa 3y agoI said "constructed algorithmically". Ie. I expect source to be at the level the engineers who built it generally worked at. It's very nice that they released their build artifacts. It's great that you can take that and make small modifications to it. That doesn't make it open source. > For most of these LLMs (not sure about this one in particular yet) the energy cost in particular of recreation is more than most individuals earn in a lifetime, and the enormous data volume is such that the only people who seriously need this should be copyright lawyers and they should be asking for it to be delivered by station wagon. All of that just sounds like reasons why it's not practical to open source it, not reasons why this release was open source.
- ben_w 3y ago> I said "constructed algorithmically". Ie. I expect source to be at the level the engineers who built it generally worked at. I could either point out that JPEG is an algorithm, or ask if you can recreate a sunset. > All of that just sounds like reasons why it's not practical to open source it No, they're reasons why the stuff you want doesn't matter. If you can actually afford to create a model of your own, you don't need to ask: the entire internet is right there. Some of it even has explicitly friendly licensing terms. An LLM with a friendly license is something you can freely integrate into other things which need friendly licensing. That's valuable all by itself.
- gary_0 3y agoI could kind of see things either way. Is this like not providing the source code, or is it like not providing the IDE, debugger, compiler, and linter that was used to write the source code? (Also, it feels a bit "looking a gift horse in the mouth" to criticize people who are giving away a cutting-edge model that can be used freely.)
- monocasa 3y ago> I could kind of see things either way. Is this like not providing the source code, or is it like not providing the IDE, debugger, compiler, and linter that was used to write the source code? Do the engineers that made this hand edit this file? Or did they have other source that they used and this is the build product? > (Also, it feels a bit "looking a gift horse in the mouth" to criticize people who are giving away a cutting-edge model that can be used freely.) Windows was free for a year. Did that make it open source?
- godelski 3y ago> Do the engineers that made this hand edit this file? Or did they have other source that they used and this is the build product? Do any open source product provide all the tools used to make software? I haven't seen the linux kernel included in any other open source product and that'd quite frankly be insane. As well as including vim/emacs, gcc, gdb, X11, etc. But I do agree that training data is more important than those things. But you need to be clear about that because people aren't understanding what you're getting at. Don't get mad, refine your communication. > Windows was free for a year. Did that make it open source? Windows didn't attach an Apache-2.0 license to it. This license makes this version of the code perpetually open source. They can change the license later, but it will not back apply to previous versions. Sorry, but this is just a terrible comparison. Free isn't what makes a thing "open source." Which let's be clear, is a fuzzy definition too.
- monocasa 3y agoWhat I'm asking for is pretty clear. The snapshot of code and data the engineers have checked into their repos (including data repositories) that were processed into this binary release. > This license makes this version of the code perpetually open source. It doesn't because they didn't release the source. There's nothing stopping me from attaching an Apache 2 license to a shared library I never give the source out to. That also would not be an open source release. There has to be actual source involved.
- godelski 3y ago> a 14GB of a dump of mainly numbers that were constructed algorithmically are not "source". I'm sorry, but what do you expect? Literally all code is "a bunch of numbers" when you get down to it. Realistically we're just talking about if the code/data is 1) able to be read through common tools and common formats and 2) can we edit, explore, and investigate it. The answer to both these questions is yes. Any parametric mathematical model is defined by its weights as well as its computational graph. They certainly provide both of these. What are we missing? The only thing that is missing here is the training data. That means of course that you could not reproduce the results were you to also have tens of thousands to millions of dollars to do so. Which if you're complaining about that then I agree, but this is very different from what you've said above. They shouldn't be providing the dataset, but they should be at least telling us what they used and how they used it. I would agree that it's not full "open source" when the the datasets are unknown and/or unavailable (for all intents and purposes, identical). The "recipe" is missing, yes, but this is very different from what you're saying. So if there's miscommunication then let's communicate better instead of getting upset at one another. Because 14G of a bunch of algorithmically constructed numbers and a few text tiles is definitely all you need to use, edit, and/or modify the work. Edit: I should also add that they don't provide any training details. This model is __difficult__ to reproduce. Not impossible, but definitely would be difficult. (within some epsilon, because models are not trained in deterministic manners, so training something in identical ways twice usually ends up with different results)
- monocasa 3y ago> I'm sorry, but what do you expect? Literally all code is "a bunch of numbers" when you get down to it. Realistically we're just talking about if the code/data is 1) able to be read through common tools and common formats and 2) can we edit, explore, and investigate it. The answer to both these questions is yes. Any parametric mathematical model is defined by its weights as well as its computational graph. They certainly provide both of these. I expect that if you call a release "open source", it's, you know, source. That their engineers used to build the release. What Mistral's engineers edit and collate as their day job. > The "recipe" is missing, yes, but this is very different from what you're saying. The "recipe" is what we generally call source. > So if there's miscommunication then let's communicate better instead of getting upset at one another. Who's getting upset here? I'm simply calling for not diluting a term. A free, permissive, binary release is great. It's just not open source. > Because 14G of a bunch of algorithmically constructed numbers and a few text tiles is definitely all you need to use, edit, and/or modify the work. Just like my Windows install ISO from when they were giving windows licenses away from free.
- cfuendev 3y agoWe should push for GPL licensing then, which AFAIK requires a source that can be built from.
- monocasa 3y agoWe just also shouldn't call releases with no source "open source". I wouldn't really have a complaint with their source being released as Apache 2. I just don't want the term "open source" diluted to including just a release of build artifacts.