3 ms·
If you take a bunch of books. Read then. Then write your own story with hints from each of the books, then who owns the copyright? Yes if you take a single work
by chrismcb 2y ago
If you take a bunch of books. Read then. Then write your own story with hints from each of the books, then who owns the copyright? Yes if you take a single work and replace every word with a synonym, that works be infringement. But that isn't what an ai does
- martin-t 2y agoLet me reiterate, there is no AI, there are only statistical models. It doesn't matter if you make the math too complex for a human to make sense of or not, it's just a remix of existing copyrighted material. Your argument forgets why copyright even exists, it's some a law of nature, it's a rule created to protect people who invest time and effort into creating. AI even if it existed would not be a person, it would be code created by a person (or more likely a corporation) to serve its goals.
- illiac786 2y agoI disagree, it does matter if the math is indistinguishable from human transformative work. What matters is the result, not the process. It’s hard to swallow with AI, but there is no other solution in my view. If we start using the process as a criteria for copyrighting, then we have chaos. Indeed, you cannot prove which process was used after the fact, especially not with AI. And the reverse is particularly problematic: anyone could then allege some human work is in fact AI and hence cannot be copyrighted. We’re breaking the copyright system if we focus on the process. Another example: if you ask an AI to create something “in the style of xxx”, then the result may be something that is infringing copyright, but so would be humans work producing tbd same output. In the end, what matters is the output, not the fact that it’s math or a human. We’re also just a bunch of atoms in the end, one could argue, very similar to a very complex mathematical model…
- martin-t 2y ago> Indeed, you cannot prove which process was used after the fact, especially not with AI. Citation needed. > anyone could then allege some human work is in fact AI and hence cannot be copyrighted I didn't say that. I said the copyright should belong to the original author. If some new work is provably based on old work to a substantial degree, it's plagiarism. Same as now. Indeed, the tool used does not matter in this case. And we know today's LLMs produce code based on copyrighted original work because they readily admit they scraped everything available to them, regardless if it was proprietary, copyleft or public domain. Older versions literally produced entire functions copy-pasted from GPL-licensed Quake code. They "patched" that now to be less blatant (mix more) but it does not make it right or legal. --- Now imagine a rogue employee of Google trains an LLM on all of Google's proprietary code (and only Google's code) and releases it publicly under, say, Apache 2.0. Can you imagine Google saying it's OK and not suing? What if that person also got his friends from Microsoft, Oracle, Apple and Amazon to pool their companies' source codes? Would that be enough mixing? Clearly the LLM would only be regurgitating code that belongs to one of these companies. What if instead somebody scraped only AGPL code and released the model under Apache 2.0? Clearly whatever comes out of the model is in its entirety based on work that is licensed under AGPL and therefore also has to be licensed under AGPL.
- Wowfunhappy 2y ago> Citation needed. How do you prove whether or not someone has used AI? > I didn't say that. I said the copyright should belong to the original author. If some new work is provably based on old work to a substantial degree, it's plagiarism. Same as now. Indeed, the tool used does not matter in this case. I am a human. I have been shaped by the world around me. Everything I have seen and read has shaped who I am today, and it has shaped the type of work which I produce myself. Yes, AI is trained on copyrighted work—but so am I! Nothing I produce is ever truly original. But it's different enough that it would not, and should not, be considered copyright infringement. Now, if I accidentally reproduced an entire function from GPL-licensed Quake code, that would be copyright infringement. And humans do make mistakes like that, when they've seen the same code many times! (Just to be clear, I am not one of those people who thinks LLMs are recreations of the human brain, I think they work differently. But we do both create output based on copyrighted input.) --- > Now imagine a rogue employee of Google trains an LLM on all of Google's proprietary code (and only Google's code) and releases it publicly under, say, Apache 2.0. Can you imagine Google saying it's OK and not suing? This is, in fact, one reason companies are always a bit worried when employees go to work for competitors. But luckily, we don't let Google say "you aren't allowed to work for anyone else after you've seen our source code."
- martin-t 2y ago> How do you prove whether or not someone has used AI? In this case it's rather easy, they admit it officially and publicly. In other cases it might be harder, we might need a whistleblower. I am sure in some cases where it happened, it'll be impossible to prove. Difficulty of proving it is not a valid reason for making something harmful legal. > But it's different enough that it would not, and should not, be considered copyright infringement. Yes, the line has to be drawn somewhere. The (typical) human brain for example has limitations how much code it can memorize and reproduce verbatim. LLMs don't have them (they're orders of magnitude higher on current hardware and models and they're essentially unlimited in principle). > But luckily, we don't let Google say "you aren't allowed to work for anyone else after you've seen our source code." But we also don't expect the new employee to build a competitor for one of google's services or internal tools in record time that can be measured in LOC per minute. An appropriately trained LLM certainly could. --- Ultimately we strayed from the main point. Copyright exists to protect authors and their intentions against parasites. If somebody's intention is to offer code for free as long as people who build on top of it also release their work for free (a slight simplification and misinterpretation of the GPL), then copyright exists to make sure that happens. That for example somebody who focuses solely on advertising (non productive zero-sum work) can't take that code and profit from it leaving the original author to eat dirt. There's also the question of income per unit of work vs "passive" income. Many builder professions get paid only as long as they're actively putting in work. We as programmers are privileged in that our product can often run without constant work and produce positive value automatically (but we only get the continuous value out of it if we own the product, not if we produced it for somebody else). Now, I believe that giving somebody fixed amount of compensation for building something that produces continuous value is fundamentally unfair, exploitative and abusive. Many people will no doubt disagree, must of the world has probably not ever thought about it by the looks of it. LLMs take this to a whole new level. They can bring in enormous amounts of money (both in subscriptions to their services and in value produced by their output). Yet the people who built the training data that made it all possible don't get _any_ compensation at all.