8 ms·
Personally, I think that the human directing the agent owns the copyright for whatever is produced, but the ability for the agent to build it in the first place
by Arcuru 5mo ago
Personally, I think that the human directing the agent owns the copyright for whatever is produced, but the ability for the agent to build it in the first place is based off of stolen IP.
I'm concerned about the copyright 'washing' this enables though, especially in OSS, and I think the right thing for OSS devs to do is to try to publish resulting code with the strongest copyleft licensing that they are comfortable with - https://jackson.dev/post/moral-ai-licensing/ https://jackson.dev/post/moral-ai-licensing/
- nadermx 5mo agoFunny how the copyright industry was able to spin copyright infringment into the pejorative "stealing". If you still have the item, what was stolen? Dowling v. United States, 473 U.S. 207 (1985): The Supreme Court ruled that the unauthorized sale of phonorecords of copyrighted musical compositions does not constitute "stolen, converted or taken by fraud" goods under the National Stolen Property Act
- themafia 5mo ago[dead]
- Neywiny 5mo agoI don't think it's unreasonable to consider it stolen potential profit, but agreed that's not how they spin it
- tensor 5mo agoI still find the idea that "learning" from code is "stealing" kind of ridiculous.
- estimator7292 5mo agoLearning, probably not. Copy/pasting at scale, yes
- vorticalbox 5mo agoIt is learning though. It’s not just copying the code. Code gets turned into tokens and then it learns the next most likely token. The issue that I see most people talk about it the scale at which is learnt. A human will learn from other people’s code but not from every persons code.
- cogman10 5mo agoThe issue is that of copyright law WRT to derivative works. Machine transformations on original works does not create a new copyright for the person that directed the machine transformation. That's why you can't pirate a bunch of media by simply adding a red pixel to the righthand corner or by color shifting the video. Copyright law is very clear that if a machine does it, the original copyright on the input is kept. This is why your distributed binaries are still copyrighted, because the machine transformed, very significantly, the source code into binary which maintains the copyright throughout. It would be inconsistent for the courts to suddenly decide that "actually, this specific type of machine transformation is actually innovative." I know this is generally really bad for the AI industry, so they just ignore it until a court tells them they can't anymore. And they might get away with it as I don't have faith that the courts will be consistent.
- red75prime 5mo agoShredding is a machine transformation. Does it mean that shreds retain original copyright even if the content can't be restored and the provenance can't be traced? Just an example that treating all machine transformations equally with no regard to the specifics doesn't make much sense. And the specifics of autoregressive pretraining is that it is lossy compression. Good luck finding which copyrighted materials have made it into the final weights.
- cogman10 5mo ago> Does it mean that shreds retain original copyright even if the content can't be restored? Yup, it absolutely does. In fact, that's why you are still violating copyright law by using bittorrent even though each of the users is only giving out a small slice or shred of the original content. The US has a granted defense in the case of something like shredding called "Fair Use" but that doesn't mean or imply that a copyright is void simply because of a fair use claim. > And the specifics of autoregressive pretraining is that it is lossy compression. That doesn't matter. Why would it? If I take a FLAC recording and change it to an MP3. The fact that it was a lossy transform doesn't suddenly give me the legal right to distribute the MP3. > Good luck finding which copyrighted materials have made it into the final weights. That's what the NYT v. OpenAI lawsuit is all about. And for earlier models they could, in fact, pull out full NYT articles which proved they made it into the final weights. Further, the NYT is currently in discovery which means OpenAI must open up to the NYT what goes into their weights. A move that, if OpenAI loses, other litigants can also use because there's a real good shot that OpenAI also included their works in the dataset.
- pydry 5mo agoIf you can set a copyright trap and an LLM reproduces it I think it's pretty clear cut that it's more than just "learning". I have seen LLMs do all sorts of crap which was clearly reproduction of training material. This is also why people are most impressed with how much better it is at reproducing boilerplate rather than, say, imaginative new ideas.
- jakeydus 5mo agoRemember last year (?) when one of the major AIs produced a bit of code that included Jeff Geerling's name in a comment?
- lo_zamoyski 5mo agoIf there were the case, then imagine having to give it back!
- boh 5mo agoYes I guess there's also no such thing as stealing in torrents since the computer "learns" the data and returns it in a transcoded fashion so it's technically not a reproduction. Yes LLMs can reproduce passages from copyrighted works verbatim but that's only because it "learned" it and it's just telling you what it "knows". The mental calisthenics required to justify this stuff must be exhausting.
- idle_zealot 5mo ago> The mental calisthenics required to justify this stuff must be exhausting. It's only exhausting if you think copyright ever reasonably settled the matter of ownership of knowledge and want to morally justify an incoherent set of outcomes that they personally favor. In practice it's primarily been a tool for the powerful party in any dispute to hammer others for disrupting their business model. I think that's pretty much the only way attempting to apply ownership semantics to knowledge or information can end up.
- balamatom 5mo agoCorrect. Knowledge consists of, roughly speaking, thoughts. (a "justified true belief" - per https://plato.stanford.edu/entries/knowledge-analysis/ https://plato.stanford.edu/entries/knowledge-analysis/ - is a kind of thought) The "thinking" part of a "thinking being" - that also consists of thoughts. If your knowledges are someone's property, you are someone's property. A society where all knowledge is proprietary, is a society of ubiquitous slavery. Maybe multi-layered, maybe fractional, maybe with a smiley-face drawn on top. Doesn't matter.
- spankalee 5mo agoHumans have been known to recite entire parts from plays from memory, live in front of audiences even.
- leni536 5mo agoAnd they are legally required to license the play to do that, if it's still in copyright.
- nkrisc 5mo agoI find it more ridiculous to equate the act of a human learning with for-profit AI training without recompense to the authors of the training material.
- MagicMoonlight 5mo agoIf I “learned” your essay and handed it in, would you be happy with that?
- array_key_first 5mo agoThe "learning" isn't learning really. I mean it might be, but if you define learning to be a human endeavor than AI can't learn. It's perfectly reasonable to say it's okay for humans to do something but not okay for a computer program to do the same thing. We don't have to equate AI to humans, that's a choice and usually a bad one.
- aeon_ai 5mo agoIf one defines 'flying' to be a bird's endeavor, then humans can't fly. Now, if you'll excuse me, I need to catch a metal shuttle that chucks itself through the air on wings.
- greendestiny 5mo agoSure as a word it can be broad, as a concept in our legal system that should be much more nuanced. The relevant extension of your analogy is should birds be required to obey FAA rules? Or should plane factories be protected as nesting sites?
- nadermx 5mo agoRelevant: https://www.bluewin.ch/en/news/swiss-company-builds-airport-in-bird-sanctuary-2790388.html https://www.bluewin.ch/en/news/swiss-company-builds-airport-...
- Dylan16807 5mo agoIt's a relevant extension if you think the ability to learn from a work is a right people have that exempts them from the more general lockdown copyright would impose. If you come at it from the view of copyright being a limited set of control over some areas but not others, then if copyright doesn't block human learning it shouldn't affect anything similar either, unless a specific rule is added to make those situations be handled differently.
- tensor 5mo agoIt's also perfectly reasonable to say it's ok for a program or machine to do the same thing as a human. This has been the basis for the technological revolution since the dawn of technology.
- greendestiny 5mo agoI think that it's absurd that we've jumped to the conclusion backpropagation in neural networks should be legally treated the same as human learning. I mean I don't think think I could find a better description for following the derivatives of error in reproducing a set of works as creating a "derivative work".
- alok-g 5mo ago>> ... we've jumped to the conclusion backpropagation in neural networks should be legally treated the same as human learning. I agree. However, the reverse is also likely true, i.e., it cannot currently be denied that learning in humans is different from learning in artificial neural networks from the point of view of production of works that mix ideas/memes from several works processed/read. Surely, as the article says, copyright law talks exclusively about humans, not machines, not animals.
- greendestiny 5mo agoI understand the article - the point about 'learning' is that if the model and its outputs are a derivative works then the copyright belongs to the human creators of the works it was trained on. Edit*: Or perhaps put more pseudo legally that the created works infringe on the copyrights of the original human creators.
- alok-g 5mo agoThe part I agree to is that copyright law calls out humans specifically as the potential owners of copyright. So what you suggest seems to be the only possibility out. Calling out humans could imply that when a human reads a thousand books and then writes something basis the same but which is not a substantial copy of anything explicitly read, that human owns the copyright to the text written. Whereas, if an artificial neural network does the same (hypothetically writing the same text), it would not. The above does not follow from, imply or conclude anything about learning in artificial neural networks and humans being similar or dissimilar.
- charonn0 5mo agoIs "learning" the correct term? Or is it "plagiarism"?
- pessimizer 5mo ago"Learning" for LLMs is just as goofy and propagandistic a metaphor as "stealing" for copyright. I find it predictive of your position that you'll accept one dumb metaphor for something that we didn't need a metaphor for, but not the other. Are you for stealing and against learning? We know exactly what is happening in both cases. We can talk about that, or we can use obfuscating euphemisms that make our preferred position seem obviously true.
- thesmtsolver2 5mo ago[dead]
- blks 5mo ago“Stolen” as in “profited on IP against terms and conditions of the license”.
- NewsaHackO 5mo agoEverybody has had a complete 180 in terms of copyright protections. Before, nobody cared about downloading music, movies, TV shows, or pirating games. Now, when the copyright law is affecting them, they are gungho about protecting these billion-dollar companies' copyrights.
- jeppester 5mo agoA more logical explanation would be that there are different opinions and those who complain are usually louder.
- NewsaHackO 5mo agoYes, that's my point. They are different and contradictory opinions, which show hypocrisy.
- inexcf 5mo agoNo it is not your point. You're just arguing about a strawman that holds both of those contradictory positions.
- NewsaHackO 5mo agoYou are attempting to invoke strawman. So is your point that there is not a significant overlap between posters who think that AI companies should not be allowed to pirated use copyrighted material in their training corpus and posters who themselves pirated copyrighted material such as movies, music, games, etc.?
- Dylan16807 5mo agoYes, that is their point. Do you have evidence against it? I'm sure you can find some overlap, but I bet the vast majority is caused by people making a distinction between commercial and noncommercial piracy. I don't think there's a big cohort of piracy hypocrites.
- CWuestefeld 5mo agobut the ability for the agent to build it in the first place is based off of stolen IP. I honestly don't understand why the attitude that underlies this is so prevalent. When I write code, what I write and how I write it is informed by having read countless source code files over my education and my career. Just as I ingest all that experience to fine-tune how my later code is written, so does the LLM from the code it's seen. The immediate retort to that is that the LLM is looking at code that wasn't its to read. But I don't think that's a valid objection. Pretty much by definition, everything I've learned from has a copyright on it, and other than my own code on my own time, that copyright is owned by someone else. Much of the code that's built up my understanding has been protected by NDA, or even defense-department classifications: it wasn't mine in any way. But it still informs how I do all my future coding. By analogy: I'm also an artist, especially since my retirement. My approach to photography was influenced by Ansel Adams, and countless other artists whose works I've seen displayed in museums, or in publications and online. My current approach to painting was inspired by Bob Ross and others, and the teachers who have helped me develop. I've taken pieces of what I've seen in all their work, and all of that comes out in my photos and paintings, to varying degrees. I've taken ideas from others in code and in art, and produced something (hopefully!) different by combining those bits with my own perspective. I don't think anyone has a claim on my product because of this relationship. Likewise, I know that many of my successors have learned from my code (heck, I led teams, wrote one book about software development!). And I hope that someday my artwork has developed to the point where there's something in it that's worth someone else's attention to assimilate. I've never for a minute - even decades before the advent of LLMs - hoped or even imagined that my work would remain locked up with me, and that the ideas would follow me to the grave. As they say, we are all standing on the shoulders of giants. None of us would be able to achieve the tiniest fraction of what we have, without assimilating what has come before us. Through many layers of inheritance it's constantly being incorporated in subsequent works. In a few decades at best, I'll be dead. It probably won't be very long after that when people even forget my name. But the idea that something I've done - my work in developing software systems, or in my photography and painting - will continue to have ripples through time, inspires me and gives me hope that I'll have some tiny shred of immortality beyond my personal demise.
- jacquesm 5mo ago
- varispeed 5mo agoI find idea that the code could be copyrightable as weak. There are only so many ways to write a for loop. Similarly you can't copyright schematics (apart from exact visual representation as form of art). Code is just a schematic.
- alok-g 5mo agoNote: IANAL Copyrights already preclude short phrases for the same reason -- there are only so many ways in which short phrases could be produced. The moment a work becomes larger (large enough; AFAIK, the threshold is not precisely defined), the reasoning you applied fails to apply. The Google-Oracle lawsuit did not decide whether APIs (when large in number) are copyrightable or not.
- gspr 5mo agoLet me get this straight: Since there are only so many ways to write a for loop, you doubt that for loops are copyrightable. From this you conclude that code, in general isn't copyrightable? That's like saying "there's only so many ways to greet your neighbor, so any text that simply greets your neighbor isn't copyrightable – and therefore no text is copyrightable".
- jacquesm 5mo agoNo, that human owns the copyright on the prompt, not on the work product.
- kridsdale1 5mo agoSo I’m responsible for pushing the giant boulder at the top of the hill. The humans at the bottom who were crushed should blame the boulder, which happened to be moving.
- jacquesm 5mo agoI'm not sure what point you are trying to make.
- Aerroon 5mo agoHe's making a point about responsibility/liability. If you only get copyright for the prompt you make, but not the output, then it's like being responsible only for the prompt, but not the output. Ie he's only responsible for pushing the boulder up the hill. The fact that it rolled down from the hill and crushed someone's house "isn't his fault" (he doesn't get copyright on it).
- jacquesm 5mo agoWell, you are responsible for the consequences. Liability is simply a different thing than copyright.
- Aerroon 5mo agoThe copyright office says that you don't get copyright because you're not considered the author: https://www.copyright.gov/ai/ https://www.copyright.gov/ai/ >The Office concludes that, given current generally available technology, prompts alone do not provide sufficient human control to make users of an AI system the authors of the output. Prompts essentially function as instructions that convey unprotectible ideas. While highly detailed prompts could contain the user’s desired expressive elements, at present they do not control how the AI system processes them in generating the output. If you're not the author then why would you have to be liable for it?
- Aerroon 5mo agoCopyright isn't some natural state of being though, it's something that's granted to people by the government to "promote the progress of science and useful arts". If copyright hinders things then I think it's reasonable that exceptions would be made.
- hxtk 5mo agoThis analysis yields very different results under utilitarianism vs rule utilitarianism. Under the former, you could argue, "What I'm doing is a science or useful art, so if copyright exists to advance those things then taking a more permissive interpretation of copyright to allow my efforts to succeed is in the spirit of the law." Under the latter, you could argue, "Works get published because as a rule, researchers and artists know they have lawful recourse through copyright if the work gets used without their consent. The absence of that rule incentivizes safeguarding works by treating them as secret and each disclosure as a matter of personal trust, so the existence of that rule promotes the sciences and useful arts."
- saadn92 5mo agoI agree with this sentiment, because the person directing the agent can still direct it in a way where it'll produce a better or worse output than another person directing it.
- rectang 5mo agoCopyright laundering is an illusion. If the LLM generates output that a court decides is sufficiently derivative, and especially (but not necessarily) if the LLM was trained on the source material being infringed, then whoever redistributes the derivative output is going to be liable for copyright infringement. Creation of the LLM itself is transformative, but LLM output which infringes is not.
- 2ndorderthought 5mo agoIs it true then that if someone stole an entire code base from a vibe coded app from a non permissively licensed project and that person claimed that it was derived from an LLM and was not stolen at all that the person who stole the code is not a thief because it came from the same place? Or are they a thief because someone else copyrighted it? How do vibe coders protect themselves not knowing who else has the same derivative code or who holds the copyright first? Or can't they?
- leptons 5mo agoThe only thing a vibe coder should be able to copyright, is the prompt text they wrote. Not the output of the LLM, only the text they wrote to instruct the LLM what to do. And even that is pretty iffy, because most of it like "put a button on a page" is not copyright-able.
- amarant 5mo agoI could possibly see an argument for the owner being whoever paid for the tokens used, but honestly I think the argument for that is weaker than what you're suggesting; I'm merely playing devil's advocate here. I don't think there's even a valid argument for any other ownership model, or at least none that I can think of.
- jmaw 5mo agoI see the argument for whoever paid for the tokens. Or in the case of a free AI usage, the person who sent the prompt (or whoever they are acting on behalf of, i.e. the company they are working for at the time). The primary issue being that it's all built on stolen data in the first place.
- pc86 5mo agoEven taking the least generous interpretation of what LLMs do and saying they're just "copy/pasting others' code" it's still not stealing because the original still exists and presumably still makes money. The original has to be gone for theft to have occurred. In order to have a sane conversation about this we have to all agree not to lie.
- deleted 5mo ago[deleted]
- ako 5mo agoI've created my own DSL, and instruct Claude Code how to generate code for this DSL using skills. Since this is a new language, and not documented on the web nor on Github, Claude's ability is not based off of stolen IP. At best it's trained on other language concepts, just like we can train ourselves on code on GitHub. Maybe a good reason to create a new programming language?
- alok-g 5mo agoInteresting, but I still do not think this is as easy. The AI model is still trained on some existing works, and it is generating code in the new DSL or programming language still based some higher level ideas and expressions it has consumed during training. You have added just one more level of indirection. The output cannot anymore be verbatim copy of some existing work or non-short snippets, however, the output may still carry "expression" that are substantially similar to something pre-existing. Note: IANAL. The above is just from my current understanding.
- jmyeet 5mo agoYou can think that's how it should be. But that's not necessarily how it is. I'm reminded of the famous monkey selfie copyright dispute [1]. A photographer set up a camera and gave it to a monkey but after a legal dispute, courts decided nobody owned the copyright. I can totally see this applying here as well. Now this doesn't resolve the issue of AIs being trained on copyrighted works it had no rights to. The counterargument is that this is a derivative or transformative work but I don't believe that's settled law at all. [1]: https://en.wikipedia.org/wiki/Monkey_selfie_copyright_dispute https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
- KallDrexx 5mo agoDo you think that human directing the agent owns copyright for any legal reason? The case Community for Creative Non Violence Vs Reid (https://en.wikipedia.org/wiki/Community_for_Creative_Non-Violence_v._Reid https://en.wikipedia.org/wiki/Community_for_Creative_Non-Vio...) solidifies a supreme court opinion that someone contracting a work and directing an author does not grant authorship to the commissioner of the work, it grants authorship to the person actually doing the work. The author can grant authorship and copyright to the commissioner with a contract, but the monkey picture (and others) have solidified that only humans can be granted copyright. Since LLMs aren't human they can't hold copyright, and if the LLM doesn't have legal copyright then they don't have legal rights to assign copyright to you.
- marcus_holmes 5mo agoInteresting, though, that ownership of the code can still be transferred to the employer. So it's in the public domain (because not human authored) but owned by the employer (because the human and/or LLM was employed by the employer)? I don't really understand how this works.
- alok-g 5mo agoNote: IANAL I think what this means is that the employee may not be the copyright owner for multiple reasons, which are possibly applicable simultaneously. It does not imply that the employer owns copyright over the work that is in public domain, which would be a contradiction.
- marcus_holmes 5mo agoyeah, that makes sense
- p_l 5mo agoCopyright works on derivative rules - is the component of the work unmistakenly derived from another copyrighted work. Under at least EU AI Act, any work done by AI is not granted copyright. But it does not mean copyright does not apply, it means the amount of work credited to AI is set at 0% (simplification). A human working off another's work unless it's perfect copy will have "credit" for changes that are judged creative/transformative, meaning a human plagiarizing something still can claim to have some degree of authorship. An AI won't. In a sense, the copyright status of final work is a sort of "sum with dilution" were each work involved adds to claims, but AI's output is set at 0 - the prompt or further rework by human is not. As for employer, details vary but generally "work for hire" rules and contracts do reassignment of material rights (in EU and some other places you can not reassign moral rights which are a different thing).
- jongjong 5mo agoThis interpretation makes sense. I think even the 'fair use' clause in the US doesn't protect LLMs. One argument I've heard often is that LLMs synthesize their training set to produce novel output in the same way as a human would... That may be the case, but legally an LLM isn't a human. You can't look at the output of an LLM and say that it's 'fair use' with respect to its training set; it hasn't been established that AI has the same 'fair use' right as a human does; it's already pushing it that companies have this right (let alone an AI agent); anyway, that's just one problem... Also, this is ignoring the fact that the researchers who compiled the training set COPIED the original copyrighted data in order to produce that training set. They either copied the entire work into the training set or they fed the entire work directly into the LLM; in either case; at some point, the entire work was copied verbatim into the LLM's input layer before it was ingested by the AI. The researchers copied the copyrighted content without permission. Also, when it comes to code, the case is even more damning because the vast majority of the code which LLMs are trained on was not only copyright but subject to an MIT license (at best) and even the MIT license, which is the most permissive license in existence, still says clearly: "Permission is hereby granted, free of charge, to any person obtaining a copy of this software" The word 'person' is used very intentionally here. I think there should be several kinds of AI taxes which should be distributed to all copyright holders. There should be a tax to go to writers (and book authors), a tax to go to open source developers and a tax for the general population to distribute as UBI to account for small-form content like comments and photography... People invested a lot of time building their entire careers around the assumption of copyright protection; so for it to be violated on such a scale would be a massive betrayal.
- amelius 5mo agoI wonder what OSS licenses would have looked like if we saw all of this coming.
- cess11 5mo agoThe LLM is just a database. It's like saying 'I own the copyright to what comes out of an API because I crafted the query' or 'I own the copyright to the responses I get from the bots on the Starship Titanic because I crafted the message they respond to'.
- dredmorbius 5mo agoThat's not what's been established to date in US caselaw: THALER v. PERLMUTTER (2023). "[T]his case presents only the question of whether a work generated autonomously by a computer system is eligible for copyright. In the absence of any human involvement in the creation of the work, the clear and straightforward answer is the one given by the Register: No." <https://caselaw.findlaw.com/court/us-dis-crt-dis-col/114916944.html https://caselaw.findlaw.com/court/us-dis-crt-dis-col/1149169...>.