5 ms·
> The output of a machine simply does not qualify for copyright protection – it is in the public domain. The machine, such as it is, is generally not acting on
by jclulow 2y ago
> The output of a machine simply does not qualify for copyright protection – it is in the public domain.
The machine, such as it is, is generally not acting on its own. A person operates the machine, and presumably is on the hook for infringement on some level.
Consider: what if one directs the machine to reproduce a specific body of code and it ostensibly does so. Was there copying?
What if I have a person read out that body of code and I type exactly what I hear? I used a machine to produce the resultant text file, but it's pretty clear that copying has occurred then.
FWIW, I'm not a copyright maximalist, but I don't think you can win a conflict by abstaining from playing the game the other team is playing. That is: the companies producing, and drawing incredible and exclusive economic value from, industrial scale plagiarism machinery are hardly going to stop ruthlessly enforcing copyright on their own proprietary software. It would seem best, if it's the goal, to get laws changed in advance of simply declaring, lopsidedly, that this is all fine.
- dtech 2y agoThe machine argument also rings hollow to me. This same argument could be made for a scanner + printer that does some transformation - changes colors a bit randomly or something - I'd be having a hard time convincing a court the resulting image is now copyright free. Obviously LLMs are much more advanced than this, but in the basis it's still a machine that takes its input data and applies specified transformations with randomness.
- lifeformed 2y agoThe machine argument makes sense to me. It shifts the blame from the machine creators to the machine users. A copying machine creator is not responsible for misuse of the machine, nor is it illegal to simply Xerox a copyrighted work. Distributing the copy is where the law comes into play, and it targets the person distributing it, not the machine or the manufacturer. Obviously, it seems impossible for an LLM user to verify the legality of the output, so it seems like the only conclusion is not to use it, or to only release your works under copyleft. I guess an alternate interpretation is treating them like gun manufacturers. They aren't the ones pulling the trigger, but one could argue their business and marketing practices are done negligently enough for them to carry a portion of the responsibility. I guess then one must show that the LLM creators are sufficiently negligent in preventing misuse of their product at the same scale.
- dtech 2y agoThe difference here is that Xerox isn't trying to handwave away copyright, while OpenAI et al. are explicitly defending the position in court that LLM output doesn't violate copyright, not that the LLM user is responsible for copyright violations. (I'd be hilarious for them to try to take this stance thought) I.i.r.c. they even "give you a license" to use LLM output, implying that they own the copyright. Too lazy to look this up so I might be wrong there though.
- elpocko 2y agoI've coded dozens of procedural asset/"art" generators. It's really surprising to me that, apparently, the output of my creative work (i.e. the novel algorithms) is not protected by copyright.
- kimixa 2y agoI've coded a JPEG compressor. It's clear the output of the machine isn't exactly the same of the input, being lossy, so I guess the output is not protected by copyright either.
- elpocko 2y agoThe input to your JPEG compressor is an image that someone/something created, and the goal of your compressor is to replicate its input as closely as possible. It doesn't create new data, it merely encodes pre-existent data. The input to my algorithms is a single number, and the goal is to create something new and distinctive that is clearly different from anything that existed before, and that wouldn't exist without the ingenuity of my algorithms.
- kimixa 2y agoReally, the "input" to current deep learning algorithms includes all the training data. Which is where the root of the issue is. There's a reason why compression comparisons tend to include the size of the decompressor in their comparisons - my "magic algorithm can compress wikipedia down to a single bye!" is less impressive with the "decompressor" contains a copy of wikipedia.
- deleted 2y ago[deleted]
- vrighter 2y agoWe go back several decades, when "computer" meant a person with a pen and paper. The output of those computers could be copyrighted... by the computer, not you.