3 ms·
If weights are copyrightable, then models should need to obtain permission before using copyrighted works to train them. You can't have it both ways.
by feanaro 4y ago
If weights are copyrightable, then models should need to obtain permission before using copyrighted works to train them. You can't have it both ways.
- ISL 4y agoEncyclopedias and journalism are copyrightable even when they include articles about copyrighted works (books, films, etc.).
- cwmma 4y agoyeah but that's because there is human creative work in the write ups of the copyrighted work, a court might liken this more to a phone book which which doesn't add any (human) creative expression on top of it.
- feanaro 4y agoNot to mention the final product of encyclopedic work and journalism does not internalize the original it is describing in its entirety, like a language model does. In some sense, a language model takes all that there is to be taken from a given resource, and incorporates it into the weights.
- microtonal 4y agoIANAL, so this is not legal advise. I consulted a legal expert a few years ago about the status of machine learning models and they said it is really unclear. Apparently, if works are transformed enough that the original is not recognizable anymore it may not violate copyright. It hinges a lot on whether the original work is reproduced, so if you could get an LLM to spit out copyrighted texts unmodified, then it would most likely be copyright violation. But I think that doesn't really happen much in practice. On the other hand, Meta can have copyright over the model through 'copyright in compilation', which protects compiled works, regardless of the copyright of the underlying material. So, I fear that it may be possible to have it both ways. But realistically, I think we'll only know for sure when this is fought out in court. Disclaimer: again I am not a lawyer, so take this with a grain of salt.
- feanaro 4y agoEven if this interpretation is correct, it only holds if the law is not changed. But the law is not immutable. My proclamation could be considered to be in terms of what ought to be, in order for society to be just and to prevent a disproportionate accumulation of power in ultra large corporations, which is detrimental to society.
- microtonal 4y agoWell that's not the way you stated. But I completely agree that that's the way it ought to be.
- A1kmm 4y agoYou could also imagine a battle over something that started with LLaMa, but fine-tuned it using a non-spare method (i.e. changed every single parameter) so it was very marginally altered in its behaviour. Even if the base model is copyrightable (possibly a big if), there is a valid question of whether a new model which essentially optimised for something else, but used the base model as a computational shortcut to make it far cheaper to solve an optimisation problem, is still protected by the copyright holder of the base model. Most of the barrier to creating large language models is the computational cost of training, not coming up with the training set data, so if fine-tuning gets around the copyright issues and allows for better FLOSS-licenced fine-tuned models, that would probably be a good thing (although maybe it will decrease the willingness of companies doing training to release models at all).