4 ms·
Copyright is just not prepared for AI. Training with copyrighted material could become "officialy legal" under copyleft terms, at least when the amount of train
by chromanoid 2y ago
Copyright is just not prepared for AI. Training with copyrighted material could become "officialy legal" under copyleft terms, at least when the amount of training material exceeds a certain threshold.
- martin-t 2y agoThe issue is treating "AI" as something special. It is just derivative work. They are called large language _models_ for a reason. They are just statistical models of existing work. If anybody seriously thought they were intelligent, they'd be arguing for giving them personhood and we'd see massive protests akin to pro-life vs pro-choice. People (well, ML companies) only use the word "intelligence" as long as it suits their marketing purposes.
- chromanoid 2y agoI agree, but it's derivative work that infringes almost indiscriminately on all publicly available cultural goods and can be very useful while doing so. That's why I think copyleft is somewhat a fitting consequence.
- martin-t 2y agoSo you think all work produced with the help of LLMs should by required to be open sourced under a copyleft-like license? That is actually an intriguing idea and at least aligned with the reasons I use AGPL for my code.
- chromanoid 2y agoHonestly, I only thought about the models, assuming that work that uses them, will also inevitably be incorporated into them. But maybe it is a good idea to actually include the produced works. It could create a nice incentive for rich corps to pay artists to create something AI free / copyleft free. If we could extract correct attribution and even licensing out of work that was produced with AI, I don't think it would help that much. I would even assume that especially in this case the rich would profit the most. They wouldn't care having to pay thousands of artists for the pixels they provided to their AI generated blockbuster movie. It would effectively exclude the poor from using AI for compliance reasons. Or even worse rich corps monopolize the training data and then they can create content practically for free, while indies cannot use AI because they would have to pay traditional prices or give the rich corps money for indirectly using their training data.
- martin-t 2y ago> pay artists to create something AI free / copyleft free I still don't think that's enough to be fair. If their work is used to produce value ad infinitum, them any one-time payment is obviously less than what they deserve. The payment should be fractional compared to the produced value. And that is very hard to do since you don't know how much money somebody made by using the model. > It would effectively exclude the poor from using AI for compliance reasons. Again, this only an issue if you're thinking in terms of one-time fixed payments.
- chromanoid 2y agoI believe you think in too short time frames. In 70 years this becomes a futile discussion. AI is a way to directly benefit from the explosion of free content that the next decades will bring. The only way to counter this in an ethical way, is to establish some kind of enforced liberation for AI models, otherwise as you say, only the rich will profit from this.
- martin-t 2y agoIt's author's life plus 70 years, if you meant that. And TBH I am mostly interested in the author's life part anyway. If it was possible to train AI models on just the public domain, them I am sure ML companies would have because it's less effort than lobbying and risking lawsuits (though I am surprised how well creators have accepted that their work is used by others to profit without any compensation, I expected way more outrage). Virtually all code relevant to training code-completion LLMs is written by people still alive or dead for way less than 70 years. We can try to come up with a better system over the next decades but those people's rights are being violated right now.
- chromanoid 2y agoI am curious why you are so critical especially when looking at code. At least for Java there are search engines to look for code that call libraries etc. Models could probably be trained on free code and then be fed with the results of these search engines on demand even through the client who calls the LLM: Client -> LLM -> Client automation API -> Code search on client machine to fill context -> Code generation w/o model that is trained on the code that was found through the code search, but merely used as context Even if they only feed code into it, that is freely given, I think the difference in quality of output would approach the current quality and better over time, especially when using RAG techniques like the above. Companies can also buy code for feeding the model after all. So beside the injustice you directly experience right now over your own code probably being fed into AI models, do you fear/despise anything more than that from LLMs?
- Kim_Bruning 2y agoEU Copyright law actually seems to have you covered already. EU Digital Single Market Directive (2019/790) Art 3 and 4 allow text and data mining. Art 3 for scientific purposes, Art 4 more in general. Now, some people argue that AI models are somehow compressed databases of the data that was crawled; but that seems patently ridiculous to me - so this should be sufficient.(at least mathematically) (IANAL) (famous last words)
- martin-t 2y agoNot ridiculous at all. If LLMs (can we please stop calling it AI?) can produce correct factual statements (for example about historical events), then the data is clearly present in the model in some (compressed) form. The only question then is if the models have some kind of additional value ("intelligence") beyond being compressed databases. My take is that either no, or the burden of proof is on those making the claim. Until they prove it, they are just databases and therefore derivative work of their input and their output is also derivative work.
- Kim_Bruning 2y agoI think you're positing a false binary. * LLM's aren't databases, you can’t query them for exact stored records, and they can’t reconstruct (most of) their training data. * But they also don’t reason or understand exactly like humans do either. They're something else: to wit, Transformer models.
- martin-t 2y agoYou have a point but 1) I don't think being able to query them and reconstruct input 1:1 are requirements. If i build a shitty db with a buggy query language that retrieves incomplete data and occasionally mixes in data i didn't ask for, then it's still a db, just a shitty one. If i populate it with copyrighted material and put it online, whether I am gonna get sued is likely based on how shitty it is, if it's good enough that people can get enough value from it that they don't buy the original works, then the original authors are not gonna be pleased. 2) Yes, comparisons to humans are not always useful though I'd say they don't reason or understand at all. Either way the discussion should be about justice and fairness. The fact is LLMs are trained on data which took human work and effort to create. LLMs would not be possible without this data (or ML companies would train on just the public domain and avoid the risk of a massive lawsuit). The people who created the original training data deserve a fair share of the value produced by using their work. So the real question to me is how much so they deserve?