4 ms·
More concretely, an ML mode is only "open source" if both the training code and the training data are freely available. Otherwise, it's not possible for the co
by cbarrick 3y ago
More concretely, an ML mode is only "open source" if both the training code and the training data are freely available.
Otherwise, it's not possible for the community to reproduce the model.
- 3cats-in-a-coat 3y agoThat's a bit like saying FOSS is open source only if a copy of the programmers is applied with the code. Everything that has a source, has another source that has produced that source. The algorithms behind creating LLMs are all published papers for all to read, the libraries (like TensorFlow) are themselves FOSS projects, and the data... is the open web for the most part. The Wikipedia dump alone is more than enough to get a very decent LLM shaped up. How an LLM is produced IS NO SECRET. It's just that to produce it you need millions (or for the more sophisticated ones: billions) in data center fees / power / GPU to train the model. So if the training scripts were included, you still can't make a LLama model yourself at home.
- tomComb 3y ago> That's a bit like saying FOSS is open source only if a copy of the programmers is applied with the code. Wait, what? You can build the FOSS app entirely from the source. You don’t need a copy of the programmer.
- JieJie 3y agoJust as you say, you do not need Llama's training data in order to build Llama from source, nor even to fine tune it. You don't need a copy of the training data. (edit to spell Llama correctly)
- KaiserPro 3y agoNo, you need llama's binary, you can;t know what its trained on because thats a secret or at least obfuscated. The currently LLM status is more akin to opensource plugins to some propriety system, like game modding.
- est31 3y agoYou can fine tune it, which gets you far, don't get me wrong. But you cannot study how major model changes to Llama or changes during the training process behave when you apply those changes from scratch.
- treyd 3y agoThe reasoning is that the model weights are a lot more like a build artifact that can't be easily scrutinized. Demanding training data to be public is analogous to demanding source code. Sources like Wikipedia are just one kind of input.
- 3cats-in-a-coat 3y agoAnd my reasoning was that everything is an artifact of something else. The model weights can't be scrutinized? Well no one can scrutinize them. Even those who made the model. And because you don't have the hardware to reproduce Llama anyway, even if they gave you the code to build it, you can't verify they used that code to build it. And if you have the hardware and data, you probably still wouldn't spend MILLIONS to end up... with the same exact model they gave you in the first place. Do you understand how meaningless this entire "Llama is not truly open" bullshit is?
- treyd 3y agoYour reasoning doesn't tackle the issue at hand so no I can't understand how "meaningless" because there is does actually matter. It doesn't matter that the people who produced the model can't scrutinize the results because they know what went into it and we don't. Just like how it's not very common for corporations that ship binary firmware blobs to scrutinize them in depth, since they have the upstream source/data that went into making them. They don't need to, because that's not the point. No you wouldn't end up with the exact same model weights if you trained it yourself on the same data but you could compare how the models perform when given the same prompts and to try to identify anomalous variation that could point you towards Facebook misrepresenting us about what went into the training / fine-tuning.
- 3cats-in-a-coat 3y agoSure you can end up with the same weights. Just start with the same random weights (put it in the source!) and have a deterministic algorithm for training (put it in the source!) and put all the same copyrighted data which Facebook can't just... put in the source, because they don't have the copyright. Do I need to keep going so you see how stupid the discussion is? You think like a software developer. Everything is just buncha text, a compiler and build environment. That's not like that. There's too much shit going in, and the rights of the data are absolutely not clear. But if they had to be, the model wouldn't exist. What you want it friggin' impossible.