14 ms·
Who owns the code Claude Code wrote?
- senaevren 5mo ago[dead]
- jhbadger 5mo agoThis is of course assuming you take AI-generated code unchanged. But you don't, in my experience. And that generates a new work fully copyrightable even if the original wasn't. Just like how the fad a decade or so ago of taking Tolstoy and Jane Austen works and adding new elements -- "Android Karenina" and "Sense and Sensibility and Sea Monsters" are copyrighted works even if the majority of the text in them was from public domain sources.
- conartist6 5mo agoI'm sure it's not quite that simple. Only parts the parts of those knock-off works that aren't public domain could be copyrightable. If you only own the copyright to ten lines in a 10k line codebase, then it's probably fair use for someone else to just to take the whole thing. Plus what if Anna Karenina was GPL?
- d1sxeyes 5mo agoAnna Karenina is public domain, assuming you’re talking about the original? If you translate it then maybe you could release it under GPL, but bit odd?
- conartist6 5mo agoI think you missed the "what if". It was just a point about how the constructed scenario might be different to the real scenario. Most AIs are not trained only on public-domain work.
- d1sxeyes 5mo agoI didn't, but not sure what the point is. Maybe I missed something else?
- brianwawok 5mo agoYou use humans to edit AI code? When you level up you are just using AI to write, AI to review, AI to edit, AI to test. Not a lot of steps left for meat bags.
- mathgeek 5mo agoYou're forgetting that you need coffee/tea/mate to fuel the button pushers. The Jetsons predicted this decades ago.
- gchamonlive 5mo agoAI for review is terrible, and by no fault of their own. It's our job to specify and document intention, domain and the right problems to solve, and that is just hard to do. No getting around it. That's job security for us meat bags.
- ModernMech 5mo agoAI to write - code is buggy and not what I asked for AI to review - shallow minutia and bikeshedding AI to edit - wrote duplicated functions that already existed AI to test - special casing and disabling code to pass the narrow tests it wrote AI report - "Everything looks good, ship it!"
- FartyMcFarter 5mo agoThe article addresses this explicitly: > Works predominantly generated by AI without meaningful human authorship are not eligible for copyright protection Note the word "predominantly", and the discussion that follows in the article about what the courts and the copyright office said.
- wongarsu 5mo agoSkimming over the article, it's a lot about what the copyright office said and very little about what courts said. But the opinion of the copyright office doesn't have any legal force. Regulations passed by the copyright office would be binding, but their opinions are just opinions. We will have to wait until relevant court cases reach a conclusion. And so far running litigation isn't even about that question, it's about infringing the rights of works that are in the training data
- throwatdem12311 5mo agoOk what about all the Anthropic’s engineers who say they don’t write code at all and it’s 100% AI-generated?
- Luker88 5mo agoNo such assumption is made in the article. Nor does it give a single answer. Mere prompting is still not enough for copyright, and the problem is unsolved on how much contribution a human needs to make to the generated code. In the case for generated images copyright has been assigned only to the human-modified parts. Even worse, it will be slightly different in other nations. The only one that accepts copyright for the unchanged output of a prompt is China.
- ModernMech 5mo agoHere's a question I have: if the AI generated image is of a character of which you own the IP, don't you have protections based on the character regardless of who gets copyright protections from authorship of the image?
- sarchertech 5mo agoYeah if you have a copyright on the character, the AI generated image doesn’t change that. It doesn’t give you more of less protection than you already had.
- beej71 5mo agoIANAL but this sounds more like trademark territory.
- sarchertech 5mo agoYou can also trademark a character if it’s used as a brand identifier in commerce. There are far more characters protected by copyright than trademark.
- gchamonlive 5mo ago> This is of course assuming you take AI-generated code unchanged. How much code do you need to change in order for it to be original? One line? 10%? More than 50%? That's arbitrary and quite unproductive convo to be honest.
- ninkendo 5mo ago> That's arbitrary and quite unproductive convo to be honest. Yeah but that’s what the legal system ostensibly does. Splitting fine hairs over whether a derived work is “transformative” is something lawyers and judges have been arguing and deciding for centuries. Just because it’s hard to define a bright red line, doesn’t mean the decision is arbitrary. Courts will mull over whether a dotted quarter note on the fourth bar of a melody constitutes an independent work all day long. It seems absurd, but deciding blurry lines are what courts are built to handle.
- stvltvs 5mo agoBecause at the end of the day, someone has to own the code, so some lines have to be drawn no matter how arbitrary they seem.
- gchamonlive 5mo agoEDIT: I changed my argument completely. That makes no sense because what if you refactor your code ad infinitum using AI? You spin up a working implementation, then read through the code, catalog the changes like interface, docs, code quality and patterns and delegate to the AI to write what you would. It's 100% AI code and it's 100% human code. That distinction is what's counterproductive.
- 6stringmerc 5mo agoWrong. This territory was heavily covered in music before this code concept - it has to be “transformative” in the eyes of the law. Even going in and cleaning up code or adding 10-25% new code won’t pass this threshold. Don't bother arguing with me on this, just accept reality and deal with it.
- jhbadger 5mo agoMy copy of "Sense and Sensibility and Sea Monsters" is explicitly listed as being copyrighted by Ben H. Winters in 2009 despite the majority of the words being Austen's, though. Perhaps music has different rules compared to text. I suspect Winters and his publisher have investigated the legality of this more than either of us have.
- acdha 5mo agoJane Austen died long enough ago that her works are in the public domain, so Winters did not need a license to use it. That does not mean that he gained rights to her work: if he tried to sue someone for use of anything which appeared in the original, he would lose in court because it’s easy to show that copies made before he was born had the same text. This also how they prevent people trying to extend copyright by making minor changes to an existing work: the new copyright only covers the additions. There’s a very accessible summary of the United States rules here: https://www.copyright.gov/circs/circ14.pdf https://www.copyright.gov/circs/circ14.pdf
- mzl 5mo agoIf you modify the work, that creates a derived work from whatever copyright the original works has, not a new work that is fully copyrightable. As the article says in the Tl;DR at the top the code may be contaminated by open source licenses > Agentic coding tools like Claude Code, Cursor, and Codex generate code that may be uncopyrightable, owned by your employer, or contaminated by open source licenses you cannot see
- exe34 5mo ago> This is of course assuming you take AI-generated code unchanged. But you don't, in my experience. And that generates a new work fully copyrightable even if the original wasn't. That's not how copyright works. The modified version is derivative. You can't just take the Linux kernel, make some changes, and slap a new license on it.
- jugg1es 5mo agoI want this question to have an interesting answer, but everyone knows that if this question ever goes to the courts, ownership will go to the people in charge with the money. The idea that Anthropic may not own Claude Code just because Claude wrote it is wishful thinking.
- embedding-shape 5mo agoBest part is, it's likely to have a different answer in every country, who knows what'll happen, not every country implicitly sides with the ones with the most money.
- adrianN 5mo agoDepends on where they pay their taxes generally.
- MarsIronPI 5mo agoWell, eventually it'll probably be added to the Berne Convention agreement or some such.
- LawnGnome 5mo agoThat's my feeling on the endgame too, but it'll probably be a decade before we get anywhere near it.
- conartist6 5mo agoIt's not wishful thinking, and ownership isn't a foregone conclusion. Sure the courts could mint a communist society with a few weird decisions about property rights, but this being the US do you really suppose that's likely? There's really no legal question of any kind that models aren't people and therefore cannot own property (and also cannot enter into legal contract as would be required to reassign the intellectual property they don't and can't own)
- wongarsu 5mo ago
- bko 5mo agoThis is all well and good as an intellectual exercise, but in real life none of this matters. Almost no one thinks their code is copyrightable or seriously thinks their code is a moat. I've written the same chunks of code for a number of employers as has every engineer. We've all taken chunks from stack overflow and other places without carefully considering attribution. This comes up in a few places as a kind of vindictive battle. One example is Oracle suing Google for too closely mimicking their API in Android. Here is an example: > private static void rangeCheck(int arrayLen, int fromIndex, int toIndex) { if (fromIndex > toIndex) throw new IllegalArgumentException("fromIndex(" + fromIndex + ") > toIndex(" + toIndex + ")"); if (fromIndex < 0) throw new ArrayIndexOutOfBoundsException(fromIndex); if (toIndex > arrayLen) throw new ArrayIndexOutOfBoundsException(toIndex); } And it was deemed fair use by the Supreme Court. Other times high frequency hedge funds sued exiting employees, sometimes successfully. In America, anyone can sue you for any reason, so sure, you'll have Ellison take a feud up with Page and Brin all the way up to the Supreme Court. In 99.9% of instances none of this matter. Sure there's the technical letter of the law but in practice, and especially now, none of this matters. https://www.supremecourt.gov/opinions/20pdf/18-956_d18f.pdf https://www.supremecourt.gov/opinions/20pdf/18-956_d18f.pdf
- croes 5mo ago> Almost no one thinks their code is copyrightable Then why does reverse engineered code need to be a clean room implementation? Ask any emulator developer or the developers of ReactOS https://reactos.org/forum/viewtopic.php?t=21740 https://reactos.org/forum/viewtopic.php?t=21740
- deleted 5mo ago[deleted]
- freedomben 5mo ago> Almost no one thinks their code is copyrightable or seriously thinks their code is a moat. You'd be surprised! Among non-software management types, they often think of the code as extremely valuable IP and a trade secret. I'm a CTO and I've made comments before to non/less technical peers about how the code (generally speaking) isn't that big of a secret, and I routinely get shocked expressions. In one case the company almost passed on a big contract because it required disclosure of the source code (with an NDA). When I told them that was a silly reason and explained why, they got it, but the old way of thinking still permeates and is a hard habit to break. Edit: Fixed errant copy pasta error. Glad that wasn't a password :-)
- DeathArrow 5mo agoI have a wood cutting machine and some wood. Who owns the timber?
- bell-cot 5mo agoSadly, IP "ownership" and copyright law are vastly more complex than ownership of physical stuff. Or were you planning to reproduce the (say) Ford Motor Company's trademarked symbol in wood? If so, you're right back in the stinkin' swamp.
- croes 5mo agoWhat is the wood in your example? This is like a machine you ask for timber and you get timber but you didn’t need to provide any wood
- deleted 5mo ago[deleted]
- skadge 5mo agoThis seems to be grounded in US law. Does anyone know if the same rules would apply in eg EU law?
- senaevren 5mo ago[dead]
- zvr 5mo agoMost of this is based on Copyright legal framework, which is surprisingly homogeneous around the world. The discussions about ownership of AI-generated material are exactly the same in EU.
- nairboon 5mo agoCopyright law kind of transcends national borders by certain international treaties like the Berne Convention. Which is why the US copyright holders could enforce their "woulnd't steal a car" threats in Europe.
- smashed 5mo agoThe "if you generated the code at work using company tools, it's owned by your employer" affirmation in the article makes no sense to me? If computer generated code is not copyrightable, ownership cannot be reassigned either.
- conartist6 5mo agoIt is copyrightable. A *human* can copyright code they wrote.
- smashed 5mo agoI meant in the sense that the "tool" is an LLM and the "work" was vibe coded. If vibe coded work is not copyrightable, it cannot be reassigned to the employer and become copyright protected.
- conartist6 5mo agocorrect
- senaevren 5mo agoThis is the sharpest point in the thread. You are right if the output has no copyright to begin with, there is nothing to assign. The employer's contractual claim over purely AI-generated code is not a copyright claim, it is a trade secret and confidentiality claim. Those are weaker protections: they require the information to remain secret, they do not survive disclosure, and they cannot be enforced against independent creation of the same code. Most IP assignment clauses in employment contracts were not drafted with this scenario in mind and may be claiming rights that do not legally exist.
- croes 5mo agoHow is it for human developers now if the company tool is a cloud tool and not running on company servers?
- deleted 5mo ago[deleted]
- padmabushan 5mo agoFirst answer who owns the model built with public data
- senaevren 5mo agoThe model ownership question and the output ownership question run on separate legal tracks and the piece focuses on the second deliberately. On the first: the model weights are owned by Anthropic under work-for-hire from their engineers regardless of what the training data contained. Training data copyright infringement is a separate tort claim against Anthropic, not a basis for anyone else to claim ownership of the model. The Bartz settlement resolved the pirated books claim without disturbing Anthropic's ownership of the weights. Owning the training data does not give you ownership of the model trained on it, any more than owning the paint gives you ownership of the painting.
- _flux 5mo agoI think it should be pretty clear that if you provided the tool the specification for the code you want, you have already provided creative input. After all, is this not what happens with compilers as well? LLM agents are just quite advanced compilers that don't require the specification to be as detailed as with traditional compilers.
- hypercube33 5mo agoTo me this is like asking who owns the binary files a compiler generates.
- yodon 5mo ago>it should be pretty clear that if you provided the tool the specification for the code you want, you have already provided creative input. If you provided a human contractor with the specifications for the code you want, the courts have repeatedly made clear you have not provided the creative input from a copyright perspective, and the contractor needs to explicitly assign those rights to you if want to own the copyright on the code.
- _flux 5mo agoLet's say we didn't have assemblers, but instead we would have three professions: - Specifiers, who make the specification for the system - Programmers, who write C code - Machine encoders, that take that C code and write machine code for a CPU Would it be that the copyright would then belong to programmers, if no other explicit assignments would be made? --- Thinking about it, probably yes: copyright of the spec belongs to specifies, copyright of the C belong to programmers, and copyright of machine code to machine encoders. Or would it depend on the amount of optimizations the machine encoders would do, i.e. is it creative or not? And then does this relate to the task and copyrightability of C compiler output, where optimizations can sometimes surprise the developer?
- perlgeek 5mo agoIn music, you can have copyright for a composition (like, lyrics and sheet music), and then for a master record. If you sell a copy of a song, you generally have pay royalties to both copyright holders. So, in your example, the specifiers would own the specification, the programmers the C code, and machine encoders own the machine code. But the ownership wouldn't be complete. If you sell the machine code, you'd have to pay royalties to all three. If you only sold the C code, only to the specifiers and the programmers.
- mensetmanusman 5mo agoIt’s the same as photography. No photographer built the multibillion dollar supply chain for the optics train in a camera, nor did they build the city scape they are enjoying as a background, they simply set the stage and push a button.
- e12e 5mo agoSeems to gloss over other kinds of contamination, beyond GPL code. Code from pirated text books, the problem with the entire language model being trained on copyright data, and on the possibility of the training data containing various copyrighted code.
- embedding-shape 5mo ago> Code from pirated text books Anthropic "solved" this by intermingling the texts extracted from pirated books (illegal) with texts extracted from the physical books they bought and destroyed (legal), so no one can clearly say if the copyrighted material it spits out came from a legal source or not. Everyone rejoiced.
- senaevren 5mo agoThe intermingling argument is actually central to the Bartz settlement structure. The settlement required destruction of the pirated dataset specifically because commingled training data creates an unresolvable provenance problem. For deployers building on Claude, EDPB Opinion 28/2024 requires a documented assessment of the foundation model's training data legal basis before deployment. "We cannot tell which outputs came from which source" is not a satisfactory answer to a regulator running that assessment. wrote about it before here: https://legallayer.substack.com/p/i-read-every-edpb-document-on-llm https://legallayer.substack.com/p/i-read-every-edpb-document...
- e12e 5mo ago> books they bought and destroyed (legal) They're only legal if training is fair use - and even I don't think it's immediately clear what would be the legal status of verbatim regurgitation of code in copyright, or code protected by patents? AFAIK I (as a human developer) can't assume that I can go and copy code out of a text book, and then assume copyright and charge for a license to it?
- embedding-shape 5mo ago> They're only legal if training is fair use The judge seems to have said it's because they "transformed" the books (destroying them after digitalizing) in the process, that made it legal. > Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. - https://arstechnica.com/ai/2025/06/anthropic-destroyed-millions-of-print-books-to-build-its-ai-models/ https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
- daishi55 5mo agoI’m no lawyer but I feel that meta, my employer, wouldn’t be letting us go hog-wild with Claude code if they weren’t completely confident that they fully owned the outputs, whether we change it or not.
- senaevren 5mo agoMeta's confidence almost certainly rests on the employment contracts and IP assignment clauses, not on a legal theory that AI output is inherently copyrightable. The enterprise agreement with Anthropic assigns outputs to the licensee. The employment contract assigns work product to Meta. Those two documents together give Meta a defensible ownership position regardless of the authorship question. The interesting gap is for developers using personal accounts or consumer plans on side projects, where neither of those documents exists.
- beej71 5mo agoI don't understand how a company can have IP copyright rights on code that is inherently uncopyrightable (in the unlikely event scotus rules that way).
- elfly 5mo agoWorst case, meta will sue the programmer who produced infringing code. I mean if the code is not copyrighteable that does not mean anything; it's just public domain code except that meta will just use good old security by obscurity to protect it. If somehow a meta programmer vibes code, say, VVVVVV, and Terry Cavanagh recognizes it on his facebook feed and sues meta, and wins, all that will happen is that meta will take down the copy of VVVVVV, will fire and sue the engineer that vibe coded it and call it a day.
- sarchertech 5mo agoThere’s so much FOMO right now around AI that no one is thinking clearly. I wouldn’t be so confident in your company.
- bearjaws 5mo agoArticle is incredibly fear mongering. Twice in my career the owners of a company have wanted to sue competitors for stealing their "product" after poaching our staff. Each time, the lawyers came in and basically told us that suing them for copyright is suicide, will inevitably be nearly impossible to prove, and money would be better spent in many other areas. In fact, we ended up suing them (and they settled) for stealing our copyrighted clinical content, which they copied so blatantly they left our own typos and customer support phone number in it. Go ahead, try to sue over your copyrighted code, 10 years and 100M later you will end up like Google v Oracle. What if the code is even 5% different? What about elements dictated by external constraints; hardware, industry standards, common programming practices, these aren't copyrightable. Then you have merger doctrine, how many ways can we really represent the same basic functions? Same goes with the copyleft argument, "code resembling copyleft" is incredibly vague, it would need to be verbatim the code, not resembling. Then you have the history of copyleft, there have been many abuses of copyleft and only ~10 notable lawsuits. Now because AI wrote it (which makes it _even harder_ to enforce), we will see a sudden outburst of copyleft cases? I doubt it. Ultimately anyone can sue you for any reason, nothing is stopping anyone right now from suing you claiming AI stole their copyleft code.
- senaevren 5mo ago[dead]
- p0w3n3d 5mo agoThat's quite impressive approach from the companies' perspective. Let's first use claude code and then we'll think who the code belongs to. I think that the gold rush approach happening right now around me (my company EMs forcing me to work with claude as fast as possible) show really short-sight of all the management people. First - I lose my understanding of the code base by relying too much on claude code. Second - we drop all the good coding practices (like XP, code review etc.) because claude is reviewing claude's code. Third - we just take a big smelly dump on the teamwork - it's easier and cheaper to let one developer drive the whole change from backend to frontend, despite there are (or were) two different teams - one for FE, one for BE. Fourth - code commenting was passe, as the code is documentation itself... Unless... there is a problem with the context (which is). So when the people were writing the code, they would not understand the over-engineered code because of their fault. But now we make a step back for our beloved claude because it has small context... It's unfair treatment. I could go on and on. And all those cultural changes are because of money. So I dub this "goldrush", open my popcorn and see what happens next.
- bearjaws 5mo agoI rarely see #3 yield better solutions, it's usually better to collaborate as a team on requirements and gotchas, but let one person own implementation.
- p0w3n3d 5mo agoBut both backend and front-end? Do everyone have to be full stack?
- sebastianconcpt 5mo agoAlso, it's supremely easy do the wrong abstractions long term and compromise premature internal designs that will start to starve of human mental modeling, hence explaining with accountability how things work and what the plans are when an incident happens. Also, if the wrong generalizations are introduced, coded correctly and reviewed and approved by AIs, then who's even driving really?
- nicoburns 5mo ago
- palata 5mo agoOne question I have is this: if an employee produces code predominantly generated by AI, it means that it is not copyrightable. Does that mean that the employee can take that code and publish it on the Internet? Or is it still IP even if it is not copyrightable? That would feel weird: if it's in the public domain, then it's not IP, is it?
- BlackFly 5mo agoA recipe isn't copyrightable but is still protected under trade secret law. I imagine that the same would apply. I think the major difference with software copyright is that I can just decompile your binary or copy a binary and give it to other people. For SAAS companies that don't distribute binaries, I imagine they basically have the same protections against rogue employees.
- ModernMech 5mo agoPresumably company policy would be implicated here, not copyright law. Whether or not it's copyrightable, what you create using AI is work product.
- cillian64 5mo agoTo look at it another way, just because some code I work on at my job is derived from open source MIT-licensed code doesn't mean I personally have the right to distribute it if my company doesn't want me to. I'd guess this comes under some generic "confidential information" clause in the employment contract.
- palata 5mo agoHmm your example is different: if you manually write code, there is a copyright for it whether it is derived from an MIT-licence or not. If you don't own that copyright (because your employer does), then you don't have the right to distribute it because it is not your code. If you generate the same code with AI, now it does not have a copyright. If it depends on an MIT library, then the MIT library has a copyright and you have to honour the licence. But the code you produced does not have a copyright (because it was generated by an AI). And therefore nobody "owns" it. My question is: can your employer prevent you from distributing something they don't own?
- tommy29tmar 5mo agoMaybe the useful test is not “who wrote this line?” but “can you show how it went from requirement/prompt/context to diff to human review/tests?” If you can’t, ownership is only one issue. You also can’t tell what was accepted as engineering work versus just copied output.
- senaevren 5mo agoThis is actually closer to how the Copyright Office thinks about it than the article makes clear. The registration guidance that emerged from the Thaler proceedings specifically asks applicants to describe the human creative contributions and how the AI was used. A documented workflow showing requirement, architectural decision, rejection of AI output, human restructuring, and review creates a paper trail that maps directly onto what the Office looks for. The can you show how it got here test you are describing is the practical version of the legal standard.
- hackingonempty 5mo agoNobody disputes that I own the copyright in a sound recording I made just by pushing the red button on my recorder. So it is a mystery to me that copyright to any sort of human conditioned machine generation is in dispute.
- senaevren 5mo agoThe sound recording analogy breaks down at the point where the recorder makes no creative decisions. Pressing record captures what is already there. Prompting Claude generates something that did not exist, through decisions the model makes about structure, naming, pattern, and implementation. The closer analogy is hiring a session musician and telling them the key and tempo. You own the recording under work-for-hire if they signed the right contract, but the creative expression in the performance is theirs unless explicitly assigned. The button you push to start the model is not the same button as the one on the recorder.
- CamperBob2 5mo agoFourier theory says that any sound, however complex, can be synthesized by summing sines and cosines. That's what an LLM does, if you twist the metaphor enough. It synthesizes complex outputs from simpler basis functions that are, or should be, uncopyrightable. The fact that it inferred those basis functions from studying copyrighted works doesn't seem relevant. Nor does the fact that the "Fourier sums" sometimes coincide with larger fragments of works that are copyrighted. How weird would it be if that didn't happen?
- array_key_first 5mo agoOf course it's relevant. How copyright infringement happens doesn't actually matter, all that matters is that the infringement happened. If I painstakingly recreate A New Hope frame by frame, pixel by pixel, that's infringement. Even if I technically used 0 content from the original.
- CamperBob2 5mo ago
- joshka 5mo agoIf you want to go much deeper, https://www.copyright.gov/ai/ https://www.copyright.gov/ai/ is particularly good at least on the side of comprehensiveness.
- TheFirstNubian 5mo agoThe elephant in the room, of course, is what constitutes “meaningful human authorship.” However, I cannot shake off the feeling that all user interactions with these AI models are being logged. Perhaps this may turn out to be the bigger concern in a potential legal battle than code authorship.
- senaevren 5mo agoThe meaningful human authorship question is the elephant, agreed, and the regulators have deliberately refused to quantify it for exactly the reason you describe any bright line number becomes a target to game rather than a standard to meet. The logging point is sharper than it might appear. In a copyright dispute over AI-assisted code, interaction logs could cut both ways. A plaintiff trying to establish human authorship would want the logs to show substantial architectural redirection, multiple rejections of Claude output, and documented reasoning for structural decisions. A defendant challenging that authorship claim would subpoena the same logs to show verbatim acceptance of output without modification. The practical implication i guess here,that the developers who want to preserve a copyright claim over AI-assisted code should treat their prompt history as a legal document from the start. It seems all over the world the logs are the evidence. Whether they help or hurt depends entirely on what they show.
- TheFirstNubian 5mo agoThe bit about treating one’s prompt history as a legal document has really struck a nerve with me. I’ve been keeping a separate git history solely for my prompts. Initially, the goals were simple: reuse prompts, turn some into skills, etc. But in light of the insights from the article and the discussions here, I need to treat this practice as serious business.
- deleted 5mo ago[deleted]
- metalcrow 5mo ago"if Claude was trained on the LGPL-licensed codebase and its output reflects patterns learned from that code, can the output be treated as license-free? The emerging legal consensus is probably not, and assuming it can creates significant liability for anyone shipping that code commercially." Is there any citation for this "legal consensus"? I was not aware there was any evidence backed stances on this topic as of yet
- onlyrealcuzzo 5mo agoThis sounds like a problem that's pretty easy to get around. CC does not need LGPL code. There's more than enough BSD and Apache code to go around. And they can generate synthetic data that is better than LGPL for their training. It's also a problem that does not seem feasible to meaningfully enforce. It's easy to generate CC code and lie and say you didn't. It would be hard to prove that you did, especially if you took any precautions to make it even slightly difficult that you did.
- adrian_b 5mo agoUnlike GPL, BSD and Apache licenses do not claim to also cover your non-AI-generated code that only invokes the AI-generated code. However, even if the BSD/Apache/MIT licensed code can be incorporated freely in your application, you still have no right to remove the copyright notices from it and/or to claim that you own the copyright for it. Therefore, unless the AI model has been trained only on non-copyrighted public-domain code, incorporating the generated code in your application means that you have removed the copyright notices from it, which is not allowed by the original licenses. There is absolutely no doubt that using an AI coding assistant works around the copyright laws, but it is still equivalent with doing copy and paste with fragments from copyrighted works into your source code. I consider that copyright should not be applicable to program sources, at least not in its current form, so reusing parts from other programs should be fair use, but only if human programmers would be allowed to do the same.
- onlyrealcuzzo 5mo ago> However, even if the BSD/Apache/MIT licensed code can be incorporated freely in your application, you still have no right to remove the copyright notices from it and/or to claim that you own the copyright for it. I can't speak for all licenses, but I'm familiar with at least one BSD license. That's almost the entire point of it... You cannot take their literal code and call it your own. You can derive code from it and call it your own. That's what LLMs primarily do.
- qsera 5mo agoMore interesting question is "Who wants to own it"... The answer is probably "Nobody"!
- jumploops 5mo agoAt what point is liability the only "job" left for humans?
- B1FF_PSUVM 5mo agoI think it was tor.com that last year had a story where the newbie hired for the corporate HR dept ended up being the last human left after all others were replaced. Ah, here we go, courtesy of google-ml: '"Human Resources" by Adrian Tchaikovsky, published on Reactor[...] https://reactormag.com/human-resources-adrian-tchaikovsky/ https://reactormag.com/human-resources-adrian-tchaikovsky/ '
- onlyrealcuzzo 5mo agoPresumably, every company that has non-LGPL CC code in production wants to own it...
- nine_k 5mo ago"Own" as in "be responsible for". Nobody is too keen to own a pile of semi-working trash, and extensive vide-coding can produce such piles easily.
- curt15 5mo agoNot sure why this is being down voted. Outsourcing work doesn't also outsource accountability.
- qsera 5mo agoYea, that is how I meant it.
- guywithahat 5mo ago
- kouru225 5mo agoIMO this is the greatest argument against AI as technofascism. The general public seems to believe that AI will usher in technofascism by claiming corporate ownership of AI output: the independent entrepreneur will be unable to compete against the corporations compute, every piece of data about you will be stolen and monetized by AI, and you will own nothing. But AI might in fact do the exact opposite and reverse the privatization trend that the West has been going through for the last 400 years. All of our copyright laws rely on the idea that there is a human consciousness behind the copyright. The more AI has input, the less we can claim ownership. If AI returns everything to the commons, then it results in a much more egalitarian world. Hilariously, many people, especially artists, see the return of the commons as an assault against them. They’re so captured by copyright that they assume any infringement on their copyright is inherently fascist. It’s ridiculous. Copyright is a corporations number 1 weapon when it comes to creating a moat and keeping the masses out. The original intent of copyright, in fact, was an incentive to return an idea to the commons. Experts used to hide their discoveries in order to keep them for themselves. Copyright provided an opportunity to release this knowledge and still profit. There were even several cases where it was established that those who claimed copyright could retain copyright even if the idea had been previously discovered. This created a huge incentive: release the knowledge or risk having your process copyrighted by the opposition. But that system worked because copyright could only exist for so long (14 years, doubled if they filed again.) Now copyright is a lifelong sentence at almost 100 years. The entire purpose of it has been undermined. Corporations own all your childhood and by the time you can profit off of it, it’s outdated. A world where the mainstream is primarily a commons seems to me like an egalitarian world. I’d like to live in that world.
- senaevren 5mo agoThe original bargain you describe, limited term in exchange for public disclosure, is exactly what makes the current situation strange. If AI-generated output falls into the public domain immediately, that is actually closer to the original intent of copyright than 95-year terms. The legal question is whether that outcome happens by design or by accident, and what it means for the people building products on top of AI-generated codebases right now.
- zuzululu 5mo agoI think it's pretty clear cut, whoever is paying for your agentic coding tool subscription is part of the litmus test. I use my own computer, I pay for my own subscription and I build my open source projects then the code belongs to me. If I use my company's computer, they pay for my subscription and we work on the company's projects then the code belongs to the company. In any step of the way if some copy-left or any other form of exotic open source license is violated, who pays for discovery? Is it someone in Russia who created a popular OSS library that is now owed? How will it be enforced?
- ottah 5mo agoMy opinion, copyright has mattered very little in the corporate world. Copyright is effectively meaningless with SaaS, and the compiled software ran on your machine is protected more by technical controls and EULAs. A world where copyright didn't exist for software would look nearly the same for the commercial world. Trade secrets, NDAs, and employment contracts bind workers more than copyright. The only thing that the question of copyright has real world impact is open source, but even then only for more restrictive licenses such as gpl.
- pocksuppet 5mo agoPlus companies just violate GPL everywhere billions of times with impunity (see: every phone ever) and nothing happens to them.
- thyrsus 5mo agoWhat is being licensed by the End User License Agreement (EULA) is the copyright on the code and its artefacts (executable bytes, etc.) - you can't have an EULA without having the copyright to license.
- ottah 5mo agoYou can have an eula on anything, it's a contract. You don't need copyright to enforce terms that two parties have agreed upon. The only thing copyright can do is force anyone in possession of copyrightable material to honor a eula. If you can only get software through approved channels, it's hard to avoid an eula. You would have to obtain it through the same pirated channels you have to now.
- Arcuru 5mo agoPersonally, I think that the human directing the agent owns the copyright for whatever is produced, but the ability for the agent to build it in the first place is based off of stolen IP. I'm concerned about the copyright 'washing' this enables though, especially in OSS, and I think the right thing for OSS devs to do is to try to publish resulting code with the strongest copyleft licensing that they are comfortable with - https://jackson.dev/post/moral-ai-licensing/ https://jackson.dev/post/moral-ai-licensing/
- nadermx 5mo agoFunny how the copyright industry was able to spin copyright infringment into the pejorative "stealing". If you still have the item, what was stolen? Dowling v. United States, 473 U.S. 207 (1985): The Supreme Court ruled that the unauthorized sale of phonorecords of copyrighted musical compositions does not constitute "stolen, converted or taken by fraud" goods under the National Stolen Property Act
- themafia 5mo ago[dead]
- Neywiny 5mo agoI don't think it's unreasonable to consider it stolen potential profit, but agreed that's not how they spin it
- tensor 5mo agoI still find the idea that "learning" from code is "stealing" kind of ridiculous.
- estimator7292 5mo agoLearning, probably not. Copy/pasting at scale, yes
- vorticalbox 5mo ago
- threepts 5mo agoWhoever pays for the tokens.
- pfortuny 5mo agoYou don't but nevertheless you bear the responsibility of making it public (whether in soyrce or binary form). That is what Anthropic would like.
- mlmonkey 5mo agoOn a related note, another question: who owns the paper that Claude (or OpenAI) wrote? Should such paper submissions in conferences call out the model(s) used to write the paper itself?
- jMyles 5mo agoThere is no such thing as ownership of a pattern of information. It has been an illusion, and that illusion is now fading.
- hmokiguess 5mo agoTangential but I find this an interesting parallel from a few years ago: https://www.vice.com/en/article/musicians-algorithmically-generate-every-possible-melody-release-them-to-public-domain/ https://www.vice.com/en/article/musicians-algorithmically-ge...
- briandw 5mo agoYour employer can claim your code if you use their tools to produce it. Nothing new here. This has nothing to do with AI tooling.
- Isamu 5mo agoCopyright has a lot to do with what we as a society want to protect and encourage. We want to protect an author that put the hours into creating a book, as opposed to the person creating a copy of that work. The person copying can claim they put in work too but the claim is not strong enough to override our preference to protect original authors. Part of the problem with generated works is that it is lower effort like the person copying something. It’s not an activity that demands special protection like original authorship. I believe this is a large part of the reasoning.
- torben-friis 5mo agoAI is a monster to our current copyright system - monster in the philosophical sense, that is, an example that destroys the concept. First, its creation is (claimed to be) extremely useful for society, but in order to be created it requires ignoring copyright for pretty much everything ever written. Something we kinda shrugged under the table. Then, it introduces an extreme jump down in creation effort - so if the focus is protection of effortful creation, nothing with AI use qualifies. But of course, you'd want society to benefit from effortlessness in general, spending more effort than needed in a task is the opposite of efficiency.
- 6d6b73 5mo agoLLMs are just tools we use. If I program an app in C++, do I not own the rights to the executable because my compiler wrote machine code for me?
- semiquaver 5mo ago> The US Copyright Office confirmed this in January 2025, and the Supreme Court declined to disturb it in March 2026 when it turned away the Thaler appeal. Works predominantly generated by AI without meaningful human authorship are not eligible for copyright protection, and that rule is now settled at the highest judicial level available. Misstates the law. Denial of certiorari can happen for many reasons unrelated to the merits and does not settle the issue nationwide.
- senaevren 5mo agoFair and correct. Cert denial means the Court declined to hear the case, not that it endorsed the lower court's reasoning or settled the question nationally. The DC Circuit ruling stands and the Copyright Office's position is consistent, but that is stable doctrine rather than Supreme Court-settled law. Updated the piece to reflect this distinction accurately.
- sowbug 5mo agoSince this is a tech audience... the Supreme Court uses a bounded priority queue. An unbounded queue would risk growing impractically large. There are some kinds of cases where the Court has "original jurisdiction," meaning they must hear them, but those are very rare.
- greensoap 5mo agoAlso, I don't think there is any example testing the conclusion. There is no case to point at that any of the factors they listed are sufficient to convey authorship. Would love to be pointed to a case where rejecting decisions and redirecting to a different approach was deemed human authorship. What we do know is that you can disclaim the part of the code a human didn't author. In fact, the Copyright Office requires you disclose and disclaim. If anyone out there has more factual and citable sources please share.
- senaevren 5mo agoYou are right that no court has yet ruled that a specific set of human contributions to AI-assisted work was sufficient to establish authorship. What exists is the inverse: the Copyright Office has granted partial registrations where human-authored elements were separated from AI-generated elements, as in Zarya of the Dawn, where the human-written text was protected but the Midjourney images were not. The Allen v. Perlmutter case pending in Colorado is the first direct judicial test of whether iterative prompting and editing can constitute authorship. Until that decision, the positive threshold is genuinely unknown. The piece reflects this in the calibration section at the end, though your point is worth adding to the authorship discussion more explicitly.
- whattheheckheck 5mo agoAsk chatgpt deep research citing court cases and it shows dark factory swe code are not copyrightable under current precedents. Even steering it with prompts isn't enough. The guy couldn't copyright the image he made with ai, code is no different. Maybe prompts written by humans are copyrightable. Can't wait for the Billionaires to entrench in court they can steal everything for these machines and claim it as their own and maybe even reach for anything that it helps produce. Fuck that
- gspr 5mo agoI'm still flabbergasted that people – and big, visible companies with big targets on their backs – choose to keep on using the output of LLMs without having an answer to these questions. And I'm worried that once that has been sufficiently normalized, laws and interpretations of them will adapt to whatever best suits those users. Which will mean copyrightwashing of FOSS. My only hope then is that surely if free software can be copyright-washed by the big guys, then so can the little guy copyright-wash the big guys' blockbuster movies or whatever, which might lead to some sort of reckoning.
- eolgun 5mo ago[dead]
- ikrenji 5mo agothe entire US economy rides on AI. no ruling throwing a wrench into the multi trillion engine is ever going to be permitted to happen
- everdrive 5mo agoWell I don't own anything I write while working on my company. Maybe my company and Claude can fight over who owns it.
- heikkilevanto 5mo agoOwnership is one question. IMO, a more interesting question is who is responsible when the code does some real-life damage.
- ACCount37 5mo agoNo one. The usual.
- mock-possum 5mo agoWhy should it be any different than it ever was? If a release manager checked it but didn’t catch the vulnerability, they have some culpability. If the developer shipped the code without checking it, they have some culpability too. Ultimately, if they both work under an organization that they report to, they’re responsible to that organization, which is, in turn, accountable to its customers (and investors perhaps.) LLMs really change nothing about this.
- mock-possum 5mo agoI do. I used a tool to create it. I own the things I create. Anything else is just bullshit equivocation.
- reliablereason 5mo agoThis is like asking: "Who owns the text microsoft word helped you write?" Claude code is a software tool not a legal entity.
- freejazz 5mo agoNot if claude does the writing. MS doesn't write things for you, and if it did, you would not be entitled to a copyright in whatever it wrote for you.
- reliablereason 5mo agoClaude is not a legal entity, it is a software tool that outputs text based on statistics. There is a user that used a tool to create text and that user is the legal entity responsible for the text in any legal way that matters. Anything else would be completely ridiculous given current laws in most countries. It would be as ridiculous as blaming the car in a car accident where you drove over someone.
- freejazz 5mo ago> Claude is not a legal entity And? >It would be as ridiculous as blaming the car in a car accident where you drove over someone. No more ridiculous than you posting something you know nothing about. Just because you don't get the copyright doesn't mean claude does. The fact that claude is not a legal entity has no bearing on whether or not you are entitled to a copyright for a work you did not create.
- deleted 5mo ago[deleted]
- reliablereason 5mo agoIf neither the user or the tool created and is responsible for the text, who is in your mind?
- deleted 5mo ago[deleted]
- dang 5mo agoCould you please stop posting generated comments to HN? It's not allowed here, and it looks like you've done it over 30 times already. (Of course, there's no way to be certain of this, but it's what our software thinks, and the overall pattern is pretty convincing.) See https://news.ycombinator.com/newsguidelines.html#generated https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079 https://news.ycombinator.com/item?id=47340079
- deleted 5mo ago[deleted]
- senaevren 5mo agoYou are definitely right to flag it, apologize for that. I used an AI assistant for the replies, and I will make sure not to use one going forward.
- dang 5mo agoAppreciated!
- deleted 5mo ago[deleted]
- Imustaskforhelp 5mo ago@dang, just wanted to say that it seems that the response to your statement does also seem to be AI generated. Dead-internet theory is turning real day by the day, oof.
- green_wheel 5mo agoWhy do you use an AI assistant for the replies?
- imafish 5mo agoMy guess is she wants to respond to all feedback and questions but doesn't have time to do it all by hand.
- jillesvangurp 5mo agoGood overview of the issues. I'm sure there are a few nits to pick with that. But something that is overlooked is that the world is bigger than the US and it's an absolute zoo out there in terms of copyright laws in different countries. Anything you think you might understand about this topic goes out of the window if you have international customers or provide software services outside the US. Or are not actually based there to begin with. And there are treaties between countries to consider as well. Courts tend to try to be consistent with previous rulings, interpretations, etc. When it comes to copyright, there are a few centuries of such rulings. The commonly held opinions among developers that aren't lawyers are that AI is somehow different. And of course since the law hasn't actually changed, the simple legal question then becomes "How?". And the answer to that seems to involve a lot of different notions. For example, "AIs are not people, and therefore any content produced by them isn't covered by copyright to begin with" is one of the notions brought up in the article. A lawyer might have some legal nits to pick with that one but it seems to broadly be the common interpretation. So AI's don't violate copyright by doing what they do. In the same way you can't charge a Xerox machine with copyright infringements. Or Xerox. But you could go after a person using one. And another notion is that any content distributed by a human can be infringing on somebody else's copyright and that party can try to argue their case in a court and ask for compensation. Note that that sentence doesn't involve the word AI in any way. How the infringing party creates/copies the content is actually irrelevant. Either it infringes or it doesn't. You could be using AI, a Tibetan Monk copying things by hand, trained monkeys hitting the keyboard randomly, a photo copier, or whatever. It does not really matter from a legal point of view. All that matters is that you somehow obtained a copy of an apparently copyrighted work. AI is just yet another way to create copies and not in any way special here. There are of course lots of legal fine points to make to how models are trained, how training data is handled, etc. But if you break each of those down it boils down to "this large blob of random numbers doesn't really resemble the shape or form of some copyrighted thing" and "Anthropic used dodgy means to get their hands on copies of copyrighted work". I actually received a letter inviting me to claim some money back from them recently, like many other copyright holders.
- randyrand 5mo agoNormally this solved with an employment contract: "Anything you write, the copyright is transferred to your employer"
- appz3 5mo ago[flagged]
- SchemaLoad 5mo agoIs code content though? This would seem more applicable to video where people are being deceived by AI content masking as real. Rather than internal application code.
- buzer 5mo agoHow do you exactly read that article in that way? 50(1) states that AI systems which interact directly with public must inform that they are interacting with AI system. 50(2) states that AI generated synthetic audio, image, video or text content must be marked as such. However this requirement applies to "providers" of AI systems. And according Article 1(3) that is: > ‘provider’ means a natural or legal person, public authority, agency or other body that develops an AI system or a general-purpose AI model or that has an AI system or a general-purpose AI model developed and places it on the market or puts the AI system into service under its own name or trademark, whether for payment or free of charge; So it sounds like it would apply to e.g. Anthropic via Claude Code, not to users of Claude Code. It's also unclear if this would apply to the compiled output or not.
- tick_tock_tick 5mo agoNo one really cares about EU law. Even the EU itself if it's inconvenient. Hell half their government website still don't comply with GDPR (probably almost all if you don't conveniently ignore the USA shield act). They'll fine some USA companies and we'll end up with another cookie banner at the start of every piece of software and some random hacker news user will claim that's absolutely not what the law wants or needs but everyone will settle on it even the website's of bureaucracy that made the law.
- giannicmptr1000 5mo agoyo Mama -Claude
- giancarlostoro 5mo agoDid Claude Code not start out as human input? Would it not be safe to say that a reasonable amount of it is still human input? But also, just because its mysteriously "not theirs" doesn't mean they magically have to give you the code.
- rnxrx 5mo agoThe idea that the provenance of a given tool's code inherently pollutes the material it's used with seems kind of illogical. Wouldn't it follow from this premise that any code written using open source IDEs and debugged with open source debuggers and other tooling would itself then be considered copyleft? Are works written with LibreOffice not copyrightable? There's obviously a huge issue with the legitimacy and ownership of training data being fed to LLMs. That seems like an issue between the owners of that IP and the people training the models and selling them as services more than the people using the tool. Isn't this just another flavor of SCO trying to extort money out of companies using Linux?
- vicchenai 5mo ago[dead]
- kazinator 5mo ago> Code that Claude Code or Cursor generated and you accepted without meaningful modification may not be copyrightable by anyone. Except if it happens to regurgitate a significant excerpt of some existing work, then the authors of that can assert their copyright; i.e. claim that it infringes.
- alienll 5mo agoThis is the same shape as the image cases. Zarya of the Dawn already settled it for Midjourney output: human-written elements were protected, AI-generated images were not. The character design didn't get copyright even though the human picked, prompted, and curated. Code isn't different. Prompting Claude to produce a function is closer to prompting Midjourney to produce a frame than to writing the function yourself. The reason it feels different to engineers is that we're used to thinking of the compiler as the analogy. But a compiler is deterministic — same input, same output. An LLM isn't. That's the line the Copyright Office is drawing, and image cases got there first.
- Onavo 5mo agoBut is there anything stopping a human from applying for copyright in their own name? Does the fact that somebody can recreate the prompt invalidate their claim?
- alienll 5mo agoFiling isn't the gate, registration is. Copyright Office requires you to disclose AI involvement and disclaim the AI-generated parts. Zarya of the Dawn is the example — applicant filed for the whole graphic novel, got partial registration on the human-written text, refused on the Midjourney images. The reproducibility of the prompt isn't really the test. The test is whether a human made the expressive choices.
- dang 5mo agoYour comments are getting classified by our software as LLM-generated or (more likely) LLM-edited. It's impossible to be certain, of course, but if this is the case—can you please not do this? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079 https://news.ycombinator.com/item?id=47340079. LLMs are amazing of course and we use them heavily ourselves - but not for modifying text that is to be posted to HN. Doing so leaves imprints on the language that readers are increasingly becoming allergic to, and we want HN to be a place human conversation.
- teeray 5mo agoWhat if no meaningful thought was put into the code (entirely vibe-coded slop), but it’s made for your employer? Shouldn’t the work be uncopyrightable?
- openclawclub 5mo ago[flagged]
- fsckboy 5mo agoit's well known that recipes cannot be copyrighted. But recipes still are protected intellectual property by trade secret law if they are treated as a secret by the holder of the recipe. Claude code itself is a trade secret, and it is not open source, so its own copyrightability is moot till you get your hands on a copy of it with clean hands. Recipes cannot be copyrighted because they are not expressions of human creativity. Software written by AIs are also not expressions of human creativity, so the balance is tilted in favor of AI generated copy not being copyrightable. The Supreme Court or legislation could change this, and I'd guess there will be a movement to go in that direction, but till something like that succeeded it's not so.
- Culonavirus 5mo ago> Software written by AIs are also not expressions of human creativity I mean I'm not the biggest fan of AI on the planet by any means (which I think my post history would prove, lol), but isn't prompt design and steering the AI "human creativity"? In one of my AI-assisted projects I spent like a week in unending threads of posts trying to make the AI do stuff the way I wanted, testing the output, finding a bazillion of bugs and "basic bitch" solutions, asking for more robust this and edge case that. It felt like I wrote a novel. How is that not creativity (Crayon-eater or Picasso, creativity is creativity)?
- comonoid 5mo agoI wonder when my manager "prompts" me "I want the feature X and I want it fast", is his prompt a human creativity?
- stetrain 5mo agoTo some extent yes. Your output at work is based on a combination of inputs from others in your organization, and is being paid for by your employer, so the organization owns the copyright on what you make for them. I think from this view it makes sense that an LLM is a tool, and the operator of that tool (or their employer) can own the output. The tricky part is when you squint and view an LLM with training input and prompted output as a machine that launders copyrighted input into customized output that is now copyrighted by a new owner. A machine that vacuums up film reels and splices them according to a set of instructions by the user to create a compilation of recent animated Disney movies with the Shrek soundtrack superimposed would probably not pass legal challenges if the user of the tool attempted to claim full copyright on the output.
- panavm 5mo ago[flagged]
- aakresearch 5mo agoIt seems that author unironically advises to write your commit messages like this: "Restructured Claude’s module architecture, rejected initial state management approach, rewrote error handling from scratch", to have a chance at defense in potential court hearing. I find it funny, if vindicating for my personal approach. If the expectation is to "restructure, reject, rewrite" what "AI" spits out, why use "AI" at all at this point???
- mifydev 5mo agoMissed opportunity for a tongue twister: Who coded the code Claude Code code?
- lihanghanger 5mo ago[flagged]
- jerleth 5mo ago> What to preserve: Commit messages that describe what you changed and why, not just what the AI generated. “Restructured Claude’s module architecture, rejected initial state management approach, rewrote error handling from scratch” is evidence. “Add rate limiting module” is not. > The second commit message versus the first is the difference between a defensible authorship claim and a clean “Claude wrote this” record. That makes no sense to me, as the commit message is probably LLM generated as well. (and even easier to generate as it doesn't have to compile or pass automated tests).
- openclawclub 5mo ago[flagged]
- bandrami 5mo agoThis is a big question that makes my employer nervous about using LLM-generated code, along with the even-more-unresolved question "what happens if the LLM outputs an algorithm that is protected by patent?" (particularly worrying because we know the base training included patent descriptions.) Questionable copyright can often be worked around (particularly since we don't distribute source) but infringing on a patent can destroy a company.
- raggi 5mo agoLawyers I have spoken to have stated strongly that they believe collective works doctrine will provide strong protections for most mature and sizable software. I see no mention of these considerations here.
- gorgoiler 5mo agoThree things matter when it comes to eating my breakfast sandwich: 1/ Was the pork in my sausage reared on a farm that meets agricultural standards? 2/ Was the food handled safely by the kitchen that cooked my food? 3/ Does the owner of the diner pay kitchen wages in accordance with labor law? By contrast, I have no idea what went into the models I use, what system prompts have prejudiced it, and whose IP has been exploited in pursuit of my answer. That’s being charitable, really. In practice the open secret of the AI industry is that the vast majority of training data, for want of a better word even if it is likely to be the most precise description, is stolen data.
- amelius 5mo agoProbably, yes, but the burden of proof is with us not them. I'm already glad some companies have the guts to open their models because proving it for open models is probably a lot easier than for a model behind a service.
- tngranados 5mo agoThat's a matter of changing a law, it's all up to the people and their representatives. We talk as if everything is set on stone but if there really is a will, there is a way.
- wartywhoa23 5mo agoThe proof is the $stupid-billion infrastructure built and kept up to host mousetraps armed with free cheese made of virtue signalling about doing the right thing and sharing the code with the world for free.
- devsda 5mo agoThe media industry loves to quote ridiculous numbers on lost revenue due to piracy etc. May be a rough ballpark numbers will get them to do something about this theft. Can someone put a rough estimate on potential revenue loss (direct and incidental) from training AI with industry wise breakup.
- 5mo ago
- dash2 5mo agoI wrote an R library doing some simple regressions using the GPU, with Claude. I asked it to provide the same API as lm, glm and some other base R functions. It copied their code wholesale without mentioning it to me. So, now my library is GPL… which is not a big deal in this context, but it was quite a shock.
- nicman23 5mo agoi do, all of it. sorry
- unholiness 5mo ago> Here is the legal baseline, in plain terms: This particular AI-ism really encapsulates what annoys me about some AI-isms. I don't mind the delves and the em-dashes that just give away the AI source of what otherwise might be good text. But these structural pieces just feel fundamentally not for the reader. Part of it is blatant pick-me language for the human feedback ("hey look you wanted plain language I did that") and part of it feels like it's just helping the future token stream (thinking-like tokens polluting the actual text). The not-this-but-that, the sycophancy, the symbolizing-vague-significance, they all have this flavor of serving a process that's no longer there as I now need to read it. It gives a similar sickening feeling to the one I get seeing something designed by committee.
- theteapot 5mo agoThat was a rather unhelpful TL;DR.
- tiku 5mo agoSo by this logic my auto complete function before Ai also wrote 50% of my code and is not made by me, because I didn't type it. What should matter is intent, the human that gives the orders.
- gspr 5mo ago> What should matter is intent, the human that gives the orders. I'd like to hear more nuance with regards to this line of reasoning. Can you conceive of a model that contains highly non-trivial representations of IP owned by others than yourself? Can you conceive that you might "order" the model to "produce" that IP? What happens then? Try this both for "open source code" as the IP, and "the novel I wrote", and "latest Hollywood movie". The model does not have to be a real model currently available. It's just a thought experiment. Try also to elaborate on the sliding scale between "an AI model" and "a compression system".
- pjmlp 5mo agoWell, have you actually read the license for the auto complete function? Example, https://marketplace.visualstudio.com/items/VisualStudioExptTeam.vscodeintellicode/license https://marketplace.visualstudio.com/items/VisualStudioExptT...
- fooqux 5mo agoI just did. Nothing in here says who owns the resulting text. Did I miss something?
- pjmlp 5mo ago> If you hover over a line of code in your application, coding assistance services will display code strings of supported function calls available through the coding assistance service that are also present in your current code file. Coding assistive services will retrieve snippets from publicly available open source code showing how others are using those same functions. 3. THIRD PARTY COMPONENTS. The software may include third party components with separate legal notices or governed by other agreements, as may be described in the notices file(s) accompanying the software.
- jimmypk 5mo ago[flagged]
- lofaszvanitt 5mo agoThis is a non issue, since any complex thing needs a lot of human oversight, otherwise it's nothing more than a multitentacled monstrosity.
- po1nt 5mo agoWho owns the code my keyboard wrote?
- heysoup 5mo agoClaude don't write code? The LLM writes code. Claude loops the LLM into writing consistent code. Humans loop Claude into consistently looping the LLM. Who own's the code? Who owns a potato? If the code is the produce of the LLM and that costs tokens, the owner of the code is the one who paid for the tokens. Money, time or attention, someone pays for the tokens, owns the code.
- PokemonNoGo 5mo ago>Who owns a potato? I don't get what this analogy is trying to tell me but I know nothing about potato law. Is this about the Belgian potato surplus?
- heysoup 5mo ago[flagged]
- cousinbryce 5mo agoI pay to listen to songs. Do I own them?
- pnt12 5mo agoSimilar to most entertainment: you have the right to consume, but not the right to adapt into your works and distribute them. Even consumption is usually limited to private usage: in my country, a consumer subscription is not enough to broadcast in a cafe or even a waiting room.
- pelasaco 5mo agoso as i understood GPL dont cover code written by agents?
- GhostDriftInc 5mo agoThe documentation advice is practical, but commit messages and prompt logs are self-reported. "Meaningful human authorship" needs a verifiable evidentiary chain, not attestations.
- reorder9695 5mo agoThe whole thing with GPL code seems like a mess and surely couldn't be set as actual precedent, right? It is totally infeasible for me to check every single GPL project on every code hosting platform to see if the code Claude etc produced is too similar. If a set of training data used for the model was released to check against that would be one thing, but you can't honestly expect someone to check every repo available from all time to see if a model (that you are not informed of what it was trained on and therefore could reproduce) might've reproduced code from it. That's not at all like checking the dependency chain of a dependency or anything as you can just read the licence of anything you're choosing to use. Surely the precedent would have to be that a model trained on GPL code has itself been infected by GPL, and therefore must have all source/weights released too if the assumption here is that it can have embedded the code well enough to be able to reproduce it?
- akersten 5mo ago> Surely the precedent would have to be that a model trained on GPL code has itself been infected by GPL, and therefore must have all source/weights released I don't see how this follows, unless we also agree that humans who have ever read any GPL code are themselves permanently tainted and therefore cannot produce anything that isn't influenced even slightly by said code. Is it just because we think the robot does a better job at learning than we do? It's an impossible line to draw, I agree, but I don't agree that the answer is "well then everything must be considered tainted," I say the answer is "ignore a vestigial concern of a bygone era."
- tremon 5mo agoThe robot does a better job at reproduction. I don't think there exists a definition of "learning" unambiguous enough to make the claim that it learns better than humans. Specifically, published models don't learn at all -- after the training phase, the model weights are fully static.
- cozzyd 5mo agoThere's an easy solution... release your code as GPL :) (but that doesn't protect you against GPL-incompatible copyleft licenses, I guess)
- cestith 5mo agoI find it distasteful and disturbing that copyright infringement by the people training the LLM in violation of a license is considered contamination by the licensed code. It’s not contamination. The code didn’t seep into your codebase. If the LLM was trained in such a way that portions of code long enough to be protectable then the license was violated by humans. The liability for the problem doesn’t lie on the shoulders of the contributors to the originally licensed code. It lies on the people inserting it into your codebase without following the terms of the license. The article also singles out the GPL repeatedly as a source of contamination. It doesn’t mention source-available proprietary licenses. It doesn’t mention code put online with no clear license, which according to the Bern Convention and the laws in at least the United States is automatically copyright protected with no license for use by others at all. It doesn’t talk about attribution for BSD-style or CC-SA-Attribution licenses. There’s no mention of leaked proprietary code. It just singles out GPL as some sort of unique problem. This seems quite shoddy and biased for an article by someone who’s writing about the law.
- vablings 5mo agoIt is probably fair that a huge share of code that is Foss is licensed under GPL, much larger than the share of source available proprietary licensed code
- numbsafari 5mo agoI would have assumed the opposite is true. Do you have any data to back that up?
- pessimizer 5mo agoYou would assume that there is more proprietary code available to read on the internet than GPL code? Do you have any rationale for that assumption? Basically all GPL code is available on the web and there is a vast amount of it. I barely see any current non-FOSS code on the internet, although I think it would be fair to count the big projects who have been using pseudo-OSS licenses lately as proprietary. Wouldn't a safer assumption be a ratio of 10:1 or 100:1 for lines of GPL vs. lines of "shared source?"
- adamzwasserman 5mo ago[dead]
- sidewndr46 5mo agoNote to anyone reading this: the author is actively reading the comments and updating the piece based off reported issues. As a result, no meaningful discussion will take place here.
- mohamedabdallah 5mo ago[dead]