5 ms·
The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights m
by loumf 2mo ago
The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.
The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).
- bonoboTP 2mo agoCopyright protects against reprinting or reproducing the wording and expression, not using the idea expressed in there in novel contexts.
- 20k 2mo agoAI reproduces copyrighted work exactly in many cases, so it clearly infringes copyright in this sense The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI
- bonoboTP 2mo agoThe verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph. Derivative work or transformative? It's not the same.
- hparadiz 2mo agoSuddenly it's "copyright infringement" to count the amount of times one word occurs after another word. I find this whole thing so amusing.
- 20k 2mo agoAI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement
- hparadiz 2mo agoThe courts don't agree with you and I don't either. Now what?
- 20k 2mo agoCourts have ordered AI models to remove song lyrics from their training data, they most definitely do not agree with you
- hparadiz 2mo agoderivative work & fair use. end of. Not only are you not winning this one but I'm gonna laugh at you the entire time.
- 20k 2mo agoOk, but courts haven't made those rulings yet so good luck with that
- hparadiz 2mo agoBartz v. Anthropic PBC, No. 24-cv-05417 (N.D. Cal. June 23, 2025) Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)
- 20k 2mo agoDid... you read any of these? Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for There's also these parts: > its use of pirated books to create such library does not constitute fair use. Which indicates that there are tight bounds depending on the ethics of how the content was obtained Similarly with the second one >Meta moved to dismiss plaintiffs’ cause of action for direct copyright infringement only to the extent that it was premised on a theory that the software comprising LLaMA is itself an infringing derivative work. We're talking specifically about the output of the models being infringing, not whether or not the models themselves are infringing. If you read onwards >Plaintiffs’ claim for vicarious copyright infringement failed because the complaint did not allege that any output generated by LLaMA contained protectable expression that recast, transformed or adapted the books. Without “an infringing output, there can be no vicarious infringement.” Which strongly indicates the precise opposite of what you're saying, if you actually like, read the rulings
- 20k 2mo agoAI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here
- visarga 2mo agoAs long as it is not substantially similar to training set it should be OK to reuse ideas. Ideas should not be protected by copyright or we can't create anything. We should not accept "vibe copyrights". Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.
- 20k 2mo agoUnder the law, machines are not able to create copyright, and basic human involvement is not enough to change this. Eg if you click a button saying "go", that won't create copyrightable content You can use all the ideas you want, but AI cannot because its not a person, and does not enjoy the same protection under the law. The copyright holders by and large did not agree to you using their content like this If we enable this, people won't create anything because all their work will immediately be stolen by the AI models. Copyright partially exists to promote the creation of new content, because theft disincentivises novel creation
- seanmcdirmid 2mo agoAre you sure you didn’t have RAG enabled and it wasn’t putting the paper in its context for your prompt? I find it hard to believe that anything but the most commonly published papers/code would exist directly in model weights.
- loumf 2mo agoEven if so, in older disputes (see Betamax/VHS), the possibility of verbatim reproduction didn’t stop the technology from being allowed because there were other non-reproduction benefits (time shifting). Also, the restriction isn’t on a technology that could possibly reproduce something. It is on the act of using the technology to reproduce something.
- visarga 2mo ago> The output of it is also a derivative work, A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.
- kelnos 2mo agoAre you a lawyer that has tested this in court, or is this just what you want the reality to be? As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.
- 20k 2mo agoI mean, its theoretically possible that a court might rule that if an LLM outputs an exact or lightly modified piece of copyrighted work, that it won't be copyright encumbered. We'll end up in a situation where copyright doesn't exist anymore, because you can always claim that its been laundered through an AI. This seems terribly unlikely to me, because 1:1 transformations (eg copying an image into memory) are already established to count as making a copy for legal purposes, there's strong precedent around piracy There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true
- loumf 2mo agoEven after they decide, it is ok for people to disagree and work to overturn those decisions. You could just be a person that _wants_ a different reality. We have political mechanisms to do that. I am not personally affected because I don’t mind LLMs using my code and writing to learn. I have open source under MIT and similar licenses. I didn’t foresee LLMs learning from it, but it does feel like it’s in the spirit of what I intended.
- brookst 2mo agoCan you point me to the statute that makes it ok for humans to learn from copyrighted work, but forbids AI? I’ll wait.
- loumf 2mo agoYou are saying exactly what I said, but making it seem like you don’t agree with me.