4 ms·
Maybe enforcing open sourcing the models is the best route to go. At some point everything worthwhile ever created will be processed. Models can be seen as the
by chromanoid 2y ago
Maybe enforcing open sourcing the models is the best route to go. At some point everything worthwhile ever created will be processed. Models can be seen as the processed, collective cultural output of humanity. It seems fair to me to force publishing the models in the vein of some kind of copyleft clause.
- sam_lowry_ 2y agoPublishing weights? Meh. Publishing code and data would lead to abolishing copyright.
- chromanoid 2y agoCopyright is just not prepared for AI. Training with copyrighted material could become "officialy legal" under copyleft terms, at least when the amount of training material exceeds a certain threshold.
- martin-t 2y agoThe issue is treating "AI" as something special. It is just derivative work. They are called large language _models_ for a reason. They are just statistical models of existing work. If anybody seriously thought they were intelligent, they'd be arguing for giving them personhood and we'd see massive protests akin to pro-life vs pro-choice. People (well, ML companies) only use the word "intelligence" as long as it suits their marketing purposes.
- chromanoid 2y agoI agree, but it's derivative work that infringes almost indiscriminately on all publicly available cultural goods and can be very useful while doing so. That's why I think copyleft is somewhat a fitting consequence.
- martin-t 2y agoSo you think all work produced with the help of LLMs should by required to be open sourced under a copyleft-like license? That is actually an intriguing idea and at least aligned with the reasons I use AGPL for my code.
- chromanoid 2y agoHonestly, I only thought about the models, assuming that work that uses them, will also inevitably be incorporated into them. But maybe it is a good idea to actually include the produced works. It could create a nice incentive for rich corps to pay artists to create something AI free / copyleft free. If we could extract correct attribution and even licensing out of work that was produced with AI, I don't think it would help that much. I would even assume that especially in this case the rich would profit the most. They wouldn't care having to pay thousands of artists for the pixels they provided to their AI generated blockbuster movie. It would effectively exclude the poor from using AI for compliance reasons. Or even worse rich corps monopolize the training data and then they can create content practically for free, while indies cannot use AI because they would have to pay traditional prices or give the rich corps money for indirectly using their training data.
- martin-t 2y ago> pay artists to create something AI free / copyleft free I still don't think that's enough to be fair. If their work is used to produce value ad infinitum, them any one-time payment is obviously less than what they deserve. The payment should be fractional compared to the produced value. And that is very hard to do since you don't know how much money somebody made by using the model. > It would effectively exclude the poor from using AI for compliance reasons. Again, this only an issue if you're thinking in terms of one-time fixed payments.
- chromanoid 2y agoI believe you think in too short time frames. In 70 years this becomes a futile discussion. AI is a way to directly benefit from the explosion of free content that the next decades will bring. The only way to counter this in an ethical way, is to establish some kind of enforced liberation for AI models, otherwise as you say, only the rich will profit from this.
- Kim_Bruning 2y agoEU Copyright law actually seems to have you covered already. EU Digital Single Market Directive (2019/790) Art 3 and 4 allow text and data mining. Art 3 for scientific purposes, Art 4 more in general. Now, some people argue that AI models are somehow compressed databases of the data that was crawled; but that seems patently ridiculous to me - so this should be sufficient.(at least mathematically) (IANAL) (famous last words)
- martin-t 2y agoNot ridiculous at all. If LLMs (can we please stop calling it AI?) can produce correct factual statements (for example about historical events), then the data is clearly present in the model in some (compressed) form. The only question then is if the models have some kind of additional value ("intelligence") beyond being compressed databases. My take is that either no, or the burden of proof is on those making the claim. Until they prove it, they are just databases and therefore derivative work of their input and their output is also derivative work.
- Kim_Bruning 2y agoI think you're positing a false binary. * LLM's aren't databases, you can’t query them for exact stored records, and they can’t reconstruct (most of) their training data. * But they also don’t reason or understand exactly like humans do either. They're something else: to wit, Transformer models.
- martin-t 2y agoYou have a point but 1) I don't think being able to query them and reconstruct input 1:1 are requirements. If i build a shitty db with a buggy query language that retrieves incomplete data and occasionally mixes in data i didn't ask for, then it's still a db, just a shitty one. If i populate it with copyrighted material and put it online, whether I am gonna get sued is likely based on how shitty it is, if it's good enough that people can get enough value from it that they don't buy the original works, then the original authors are not gonna be pleased. 2) Yes, comparisons to humans are not always useful though I'd say they don't reason or understand at all. Either way the discussion should be about justice and fairness. The fact is LLMs are trained on data which took human work and effort to create. LLMs would not be possible without this data (or ML companies would train on just the public domain and avoid the risk of a massive lawsuit). The people who created the original training data deserve a fair share of the value produced by using their work. So the real question to me is how much so they deserve?
- marssaxman 2y agoWell, that sounds good - what are we waiting for?
- martin-t 2y agoEven if they are small enough to be run be individuals, this still doesn't solve the issue of profiting from someone's work for free. Many of us publish our open source work under GPL or AGPL with the intention that you can profit from it but you have to give back what you built on it under the same license. LLMs allow anyone to launder copyleft code and profit without giving anything back. If people who downvote bothered to reply, they'd probably say that by being "absorbed" into LLM weights, the code served that purpose and is available for everyone to use. That forgets 2 critical points: - LLMs give no attribution. I deserve to be credited for my fractional contribution to the collective output of humanity. - LLMs are not intelligences, they don't suddenly make intellectual work redundant. Using them still requires work to integrate their output and therefore companies build (for-profit) products on top of my work without compensating me, without crediting me and without giving anyone the freedom to modify the code.
- chromanoid 2y agoI am not sure if profits are easy to gather if all professional users can rent hardware instead of renting the service. In the end the users would actually profit from the model, not the provider (who competes with other providers of the same open source model). And here I am not sure if we can see AI as some kind of cybernetic enhancement that allows to execute ideas on the shoulders of humanity instead of just a scammy way to resell already present content.
- martin-t 2y agoExactly. It's impractical to distribute compensation (and credit) fairly so it's very easy for those who profit to say "we can't", their their hands up and keep profiting. Doesn't make it right. It's just the rich getting richer by taking from everyone so little that no individual bothers to fight back.
- chromanoid 2y agoWhy do you think the rich get richer when operating an AI is a price competition and training new models gives only a short competitive advantage? I would hope that for each individual there opens an ocean of opportunities that is sustained by all humans that came before and poured their bucket of knowledge into it.