5 ms·
If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspira
by flessner 1y ago
If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another.
What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.
- nickpsecurity 1y agoThe FairTrained models claim to train with only public domain and legal works. Companies are also licensing works. This company has a lawful, foundation model: https://273ventures.com/kl3m-the-first-legal-large-language-model/ https://273ventures.com/kl3m-the-first-legal-large-language-... So, it's really the majority of companies breaking the law who will be affected. Companies using permissible and licensed works will be fine. The other companies would finally have to buy large collections of content, too. Their billions will have go to something other than GPU's.
- bilbo0s 1y agoI don't know? Not really sure a claim is good enough. I don't know that you can just go into court and say, "Trust me, I don't use copyrighted material." And I also can't see any way, other than providing training data and training an identically structured model on that data, that a company can conclusively show that they got the weights in an allegedly copyright free model from the copyright free training data a company provides.
- hobs 1y agoCivil courts work by you proving damages (at least in the USA), not by you going on fishing expeditions because they "might" have done something. So good luck finding the thing that looks exactly like your copyrighted work that's not in the corpus, if you can yeah, you might be able to prove it. At the end of the day its like a lot of business, where a liability shell game is played out, and if the chain of evidence cant be drawn quite brightly then lawsuits would be frivolous at best.
- deleted 1y ago[deleted]
- 317070 1y agoI do hope people are still innocent until proven guilty? If you did not use copyrighted materials for training, people will not be able to prove that you did, and that should be good enough.
- lelanthran 1y ago> I do hope people are still innocent until proven guilty? It's a civil matter not a criminal matter so that that doesn't apply.
- nickpsecurity 1y agoWhile the others are correct, I'm with you in the sense that I don't know if what they claim is true. I've also found others, like one in Singapore, that didn't use it on data that was as legal as news reports claimed. It might turn out to have problems. There is benefit to using them, though. For one, they've tried really hard to be legal. That sets a positive example, shows good faith if they were sued, and reduces risk for those using them (good faith on our part). Also, one can be sure that they can ditch or replace any outputs in the long term if they're ruled illegal. So, we try not to use the A.I.'s in a way where losing access to them seriously damages our business. That's the best I can offer until legal reforms happen. If training, one can train it in Singapore on material you he or she has legal access to. Their law pretty much let's you use anything for AI purposes so long as you legally can access it yourself. To further reduce the risk, they should crawl it themselves, too, taking care to avoid risky sources.
- ekidd 1y ago> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the US decides to unilaterally shut down LLMs, that just means that the rest of the world will route around us. Whether this is good or bad is another question.
- theturtletalks 1y agoChina found the perfect way to disrupt US tech, releasing open source versions of it for free or at least cheaper. Most of US tech is built on open source anyways and with the pace YC is investing in open source alternatives, it will win out in most niches. My fear is that the US tech won’t be able to compete with state sponsored open source out of China and will move to ban open source or suppress it somehow.
- ekidd 1y agoAlso, the Chinese work is legit. DeepSeek introduced a whole bag of new techniques like GRPO, and released quite a bit of good open source tooling. And Alibaba's Qwen team seems to be quite genuinely talented at "small" models, 32B parameters and below. Once you get Qwen3 properly configured, it punches well above its "weight class." I'm still running real benchmarks, but subjectively, it feels like the 32B model performs somewhere between 4o-mini and 4o on "objectively measureable" tasks. It's a little "stodgy" and formal by default, though. We'll see what it looks like when people start fine-tuning it. If the US dropped off the planet, it would maybe set LLM technology back a year.
- theturtletalks 1y agoDeepseek really changed how people think about Chinese tech. Even after new LLMs launched, Deepseek R1 and V3 hold their own on benchmarks and are significantly cheaper.
- Yizahi 1y agoThe whole point of the fair use clauses is to protect humans. Clearly we can easily say that programs are altogether exempt in favor of humans, and it would be a proper thing to do, until the first real AI is built.
- diggan 1y ago> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. End of the road for major AI companies, and hopefully something better can be created once it's declared illegal without any murky waters. There are LLMs trained on data that isn't illegally obtained, OLMo by Ai2 is one such model, that is actually open source and uses open data for training. Just because it's "very difficult" for OpenAI et al shouldn't be an argument to force them to behave ethically anyways. If they cannot survive acting legally, then so be it, sucks for them.
- nradov 1y agoThat would hardly be the end of the road. If copyright enforcement gets stricter then that will give a market advantage to the largest, best funded major AI companies like OpenAI because they can afford to simply buy licenses from copyright holders. I predict that we'll see new middlemen arise specifically to handle this licensing, much like the agencies that handle most music licensing today.
- dragonwriter 1y ago> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. You assume that getting tested means the AI trainers lose, and also thar the model architectures that have been developed can’t be retrained from scratch with public domain, owned, and purpose-licensed material. (With several AI companies having been actively pursuing deals to license content for AI training for a while now.)
- vkou 1y agoIf corporations owned human slaves and fed them copyrighted materials so that they were inspired to produce original creative output, I don't think that creative output should enjoy legal protections either. Even if slavery were not illegal. Because the obvious question would be - how can free people compete with that?
- const_cast 1y agoThe lines for humans aren't clearly drawn, but they are drawn. The main difference is that humans are humans and LLMs are computer programs. I see no reason why we should even entertain the idea of extending human rights to computer programs, and so far, nobody has been able to give me any good reasons why. Furthermore, why are we only entertaining the human rights that can be used for profit-driven purposes? Why do LLMs, for example, not have the right to free speech? Or an attorney? It seems highly unethical to grant these computer programs some protections as if they're humans but not grant them personhood. This is akin to slavery, which is something we actually have to consider. Anthropomorphization is a double-edged sword. We cannot simultaneously consider them human when convenient and then consider them programs when it's not. Or, if we want to do that, we need to form coherent argument to why, how, and when.
- EMIRELADERO 1y agoYou're thinking about it using the wrong framework IMO. It's not about the program's rights, it's about the human's rights to use the program. Not the machine's right to do something, but the human's right to do something through a machine, or make a machine do something.
- const_cast 1y agoNo, because the entire argument hinges on the fact that LLMs learn, which is like humans learning, so it's transformative. That only works if you consider learning or transformation to be something that does not rely on the human spirit. Which, actually, most people do not believe. And it's pretty difficult to argue - we don't even know how learning works for people. A lot of people just jump to LLMs learning like it's a foregone conclusion. Mm... no. You need to convince people of that. You'll find if you talk to non-tech people, they're not just going to believe you when you say that. Why isn't an LLM more akin to a database or a compression algorithm? Why is it closer to human learning? After all, humans are humans and we have the exclusive right and power to determine what is human and what isn't. And database and compression algorithms are computer programs, of the same kind as an LLM.
- aprilthird2021 1y agoIt's not the end, all these companies have "clean" datasets which they train their models on now, along with training on the previous "dirty" models. But it's been so many generations, that they don't need to worry about this copyright issue anymore
- imtringued 1y agoOr maybe they just need a license for their particular use case...