4 ms·
> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no payw
by ltadeut 3y ago
> , that data can and will get scraped if the intention is to use for a model.
How would that work from a legal perspective, though?
Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?
- EMIRELADERO 3y agoIt is OpenAI's and Microsoft's position that obtaining, using for training, and displaying data via AI is fair use.
- asylteltine 3y agoI really really really hope the copilot and other suits are successful. The idea that you can literally steal content in the name of “””AI””” and profit off it is just insane! How is it not like copyright infringement? The Warhol case is just one step behind training data. It’s basically the same idea.
- skilled 3y agoInteresting question but sadly I am in no position to answer it. I think there are probably issues to address with scraping it blindly: - Can Reddit imprint its data somehow? A watermark? - Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization? If OpenAI can't work around this, I'm not sure they would be willing to cross any lines in terms of copyright, they've already done it with ChatGPT and I am guessing rules are only going to get stricter on this topic.
- comfypotato 3y agoI think the bottom line is that Microsoft’s (and thus other for-profit AI initiatives) stance is that any and all data is fair game regardless of license or authorization. This results, in their opinion, from the fact that the AI alters the data, changes the output, and is otherwise “inspired” by the data in the same way an artist might be inspired by another without copyright infringement.
- nologic01 3y agoThis sounds very dodgy. Will somebody be checking the degree of such "data alteration" and verify that the "AI" is actually inspired rather than copying? To me this feels like its opening up the door for the elimination of copyright as any algorithmic layer interjected between scrapped data and end users could claim to be "inspired".
- comfypotato 3y agoWelcome to the discussion lol. People have already provided examples of copilot producing niche code verbatim thereby proving their intuition incorrect. It’s a whole mess that will take years to be cleaned up by new legal conventions.
- Jevon23 3y agoIf my compiler was “inspired” by leaked Windows source code and altered it into a new form then I think their opinions on the matter would be very different.
- comfypotato 3y agoNot a great example: if Apple’s code leaked, theoretically they wouldn’t include it in the training as it’s not supposed to be seen by the public. If it’s public you can be inspired by it (so their logic goes).
- dontupvoteme 3y agoThe true malicious, and probably effective, approach is to silently poison outputs if you suspect automated behaviour. These large language model things might be useful there. Or the old school NLP stuff.
- csdvrx 3y ago> Wouldn't that be a violation of Reddit's terms and liable to some legal procedure? It won't, with the LinkedIn vs HiQ precedent.
- dom96 3y agoHow does Google use Reddit's data in its models? You can access most (all?) Reddit pages without hitting Reddit at all via the "Cached" link in the search results. Does Google have a special agreement with Reddit (and all other sites?) or is it legally "fair use" to reproduce web pages that are available freely online?
- mminer237 3y agoI think that's a different legal question than LLM training, but webpage caching has been found to be fair use based on a number of factors: https://www.pinsentmasons.com/out-law/news/google-cache-does-not-breach-copyright-says-court https://www.pinsentmasons.com/out-law/news/google-cache-does...
- mminer237 3y agoReddit's terms are irrelevant. Unless Reddit requires a login to view its site (which would also prevent Google indexing), anyone can view the data without agreeing to the terms. The only question is copyright, but I find it hard to argue that LLM training is not sufficiently transformative in 99% of cases.
- icebraining 3y agoPlus Reddit doesn't own the copyright to the posts, the users do.
- deleted 3y ago[deleted]
- tedivm 3y agoIn the US at least the courts has made it clear that scrapping is legal. https://techcrunch.com/2022/04/18/web-scraping-legal-court/ https://techcrunch.com/2022/04/18/web-scraping-legal-court/
- cma 3y agoReddit don't own the copyright to it, just a license. That plus public web scraping is legal. Reproducing the data directly might violate the user's copyrights, but through an LLM it is assumed not.