3 ms·
The models were built using copyrighted works, so why can't models be built using other models?
by foo12bar 2mo ago
The models were built using copyrighted works, so why can't models be built using other models?
- breppp 2mo agoBecause model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws. In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright
- preg_match 2mo agoWhy would this be the case. Why would software output from a model magically have greater protection than the software the model trained on.
- vel0city 2mo agoLet's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP. Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen? If the model output is owned by the person prompting it and paying for the tokens, what's the problem here? If the model output is owned by the trainer of the model, that's a big nasty can of worms.
- giaour 2mo ago> I doubt these will have worse protection than software does, which has far better protections than copyright Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.
- breppp 2mo agoSoftware is protected by the DMCA, patents, licenses, EULAs, all of those aren't there for books. I doubt new laws won't be written for model outputs. Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply
- giaour 2mo agoYou may recall that the DMCA was originally written to protect music and movies. It does in fact apply to creative works. If you have ever purchased an MP3, eBook, or streaming movie, you will also be aware that you purchased a license to the underlying IP. This is also true of physical media, but the license agreement you have to accept when obtaining a digital work makes this explicit. I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article.
- vel0city 2mo agoYou do patent ideas, but the actual words written in a book describing that idea would only be protected by copyright at best. FWIW, the exact words describing the idea being patented are technically public domain; that's the whole point. You're free to go look up that patent, print it out, make whatever copies of it you want. Take any of the drawings in patents, put them on t-shirts, and sell them. No problem. Implementing the ideas those words represent is a different story. For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea.
- bigiain 2mo agoThe C in DMCA stand for Copyright. All (I think?) software licenses are underpinned and made legally enforceable by copyrights. EULAs are underpinned by licenses which are founded on copyright. Patents are the only one of those protections that are not based on copyright, and there are lots of very good arguments against at least most software patents (all software patents of the form "Do {well known and obvious thing} with a computer" should, in my opinion, be immediately revoked and potentially have every company who's enforced payments from such patents investigated for fraud).
- wasfgwp 2mo agoLLM outputs are not copyrightable. At least that’s the current established legal precedent in the US. The only question is whether the user owns the copyright without significantly transforming the output but that’s not really relevant in those specific situation. I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..
- usef- 2mo agoThey do seem to be paying for it (as per the 1.5Bil lawsuit yesterday and them now purchasing books and licensing from media companies). Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent. We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models. (Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ https://stratechery.com/2026/whos-afraid-of-chinese-models/ )
- trhway 2mo agoThe judge found their use is fair use. They are paying not for their use of the content, they are paying for using illegal copies of the content. The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled. To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.
- usef- 2mo agoFair, but isn't "illegal" access what they're talking about in OP? It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international.
- Bratmon 2mo ago> It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI). Do you have a source?
- usef- 2mo ago