3 ms·
The codebase to do the training is way less valuable than the weights for the vast majority of people. Releasing the training code would be nice, but it doesn't
by valine 2y ago
The codebase to do the training is way less valuable than the weights for the vast majority of people. Releasing the training code would be nice, but it doesn't really help anyone but Meta's direct competitors.
If you want to train on top of Llama there's absolutely nothing stopping you. Plenty of open source tools to do parameter optimization.
- diggan 2y agoNot just the training code but the training data as well, should be under a permissive license, otherwise you cannot call the project itself Open Source, which Facebook does here. > is way less valuable than the weights for the vast majority of people The same is true for most Open Source projects, most people use the distributed binaries or other artifacts from the projects, and couldn't care less about the code itself. But that doesn't warrant us changing the meaning of Open Source just because companies feel like it's free PR. > If you want to train on top of Llama there's absolutely nothing stopping you. Sure, but in order for the intent of Open Source to be true for Llama, I should be able to build this project from scratch. Say I have a farm of 100 A100's, could I reproduce the Llama model from scratch today?
- unshavedyak 2y ago> Not just the training code but the training data as well, should be under a permissive license, otherwise you cannot call the project itself Open Source, which Facebook does here. Does FB even have the capability to do that? I'd assume there's a bunch of data that's not theirs and they can't even release it. Let alone some data that they might not want to admit is in the source.
- bornfreddy 2y agoIf not, it is questionable if they should train on such data anyway. Also, that doesn't matter in this discussion - if you are unable to release the source under appropriate licence (for whatever reason), you should not call it Open Source.
- talldayo 2y agoI will steelman the idea that a tokenizer and weights are all you need for the "source" of an LLM. They are components that can be modified, redistributed and when put together, reproduce the full experience intended. If we insist upon the release of training data with Open models, you might as well kiss the idea of usable Open LLMs out the door. Most of the content in training datasets like The Pile are not licensed for redistribution in any way shape or form. It would jeopardize projects that do use transparent training data while not offering anything of value to the community compared to the training code. Republishing all training data is an absolute trap.
- enriquto 2y ago> Most of the content in training datasets like The Pile are not licensed for redistribution in any way shape or form. But distributing the weights is a "form" of distribution. You can recover many items of the dataset (most easily, the outliers) by using the weights. Just because they are codified in a non-readily accessible way, does not mean that you are not distributing them. It's scary to think that "training" is becoming a thinly veiled way to strip copyright of works.
- talldayo 2y agoThe weights are a transformed, lossy and non-complete permutation of the training material. You cannot recover most of the dataset reliably, which is what stops it from being an outright replacement for the work it's trained on. > does not mean that you are not distributing them. Except you literally aren't distributing them. It's like accusing me of pirating a movie because I sent a screenshot or a scene description to my friend. > It's scary to think that "training" is becoming a thinly veiled way to strip copyright of works. This is the way it's been for years. Google is given Fair Use for redistributing incomplete parts of copywritten text materials verbatim, since their application is transformative: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... Or Corellium, who won their case to use copywritten Apple code in novel and transformative ways: https://www.forbes.com/sites/thomasbrewster/2023/12/14/apple-and-corellium-settle-copyright-fight/ https://www.forbes.com/sites/thomasbrewster/2023/12/14/apple... Copyright has always been a limited power.
- jncfhnb 2y agoPeople don’t typically modify distributed binaries. People do typically modify model weights. They are the preferred form to modify model. Saying “build” llama is just a nonsense comparison to traditional compiled software. “Building llama” is more akin to taking the raw weights as text and putting them into a nice pickle file. Or loading it into an inference engine. Demanding that you have everything needed to recreate the weights from scratch is like arguing an application cannot be open source unless it also includes the user testing history and design documents. And of course some idiots don’t understand what a pickled weights file is and claim it’s as useless as a distributed binary if you want to modify the program just because it is technically compiled; not understanding that the point of the pickled file is “convenience” and that it unpacks back to the original form. Like arguing open source software can’t be distributed in zip files. > Say I have a farm of 100 A100's, could I reproduce the Llama model from scratch today? Say you have a piece of paper. Can you reproduce `print(“hello world”)` from scratch?