5 ms·
So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI sys
by badcppdev 3y ago
So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something.
A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe)
But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copyright of older material.
- limaoscarjuliet 3y agoWe learned to make a distinction based solely on "mechanisation" aspect: - email vs spam, - phone call vs robo-call, - hand drawn vs a copy machine (to some extent), - etc.
- Pet_Ant 3y ago> So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. If we don't take this approach then there will be a series of very lame legal loop-holes with putting mechanical Turks [1] in the process. Or just end up with very "I know it when I see it" legislation. So for both practical and philosophical grounds I do support this. [1] https://en.wikipedia.org/wiki/Amazon_Mechanical_Turk https://en.wikipedia.org/wiki/Amazon_Mechanical_Turk
- JKCalhoun 3y agoAnd "clean room" AI training other AI. The end is inevitable, protecting copyright for AI training is probably a lost cause.
- vidarh 3y agoYou don't even need that. Even "just" OpenAI has a valuation large enough to just buy several major publishers and data brokers to secure access to data if they need to license. And while I expect NYT imagines that their archive is really valuable for training, they're just not that special in the sense that while they may have broken more stories on average than many others, and have had influential op eds etc., the ones that matters will have been cited and referenced and written about elsewhere - the irony is that by virtue of being so well known, their historically most important content is also less unique in terms of the accessibility of the information in it. So while I'm sure OpenAI would love their archives, I'm also sure that if OpenAI and others have to license content and NYT end up being "difficult", OpenAI will just license content from (or buy) a suitably diverse portfolio of other papers instead. In other words, beyond producing outright synthetic data, if AI companies are prevented from training on data they don't have a license to, the net effect will just be a scramble to buy licenses and/or buy companies that can provide sources of content, and the price for that content will be a lot lower than some of the people pursuing these copyright claims imagine. In the end, if we go that route, all we'll have achieved as a society is creating massive moats protecting the companies already big enough to buy access to a broad enough set of content and made open models harder.
- miki123211 3y agoThis will also enable new and unique business models for social network. You won't have to rely on ad targeting and user tracking any more, if you encourage people to make great content, you can make money by licensing that content to advertisers.
- jncfhnb 3y agoIt is not a problem if people can mechanism copyright infringements. If the violation is material enough to matter in a public market then we can sue.
- night-rider 3y agoGenerative AI tweaks the original material such that it's modified beyond recognition, and non verbatim. A bit like how artists tweak other artist's work, put their own spin on it, and then pass it off as 'original'.
- redcobra762 3y agoIf you ask it to sure, but there are examples of ChatGPT reproducing whole NYT articles, with not nearly enough alterations to constitute a new work.
- jondwillis 3y agoWhat were the prompts? A whole article would barely fit in the output tokens, typically, yeah?
- redcobra762 3y agohttps://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... Exhibit J (phone won’t cooperate with paste).
- Xelynega 3y agoAn alternative viewpoint: Generative AI algorithmically processes the inputs such that they can be roughly recreated after decoding. A bit like how jpeg encodes images "beyond recognition"(aka lossy).
- pclmulqdq 3y agoExcept when it doesn't. See the NYT's lawsuit here - long runs of the training data are spit out verbatim.
- rvz 3y agoOr produces the watermark from the copyright holder which even Getty can see it is theirs [0], implying that Stability AI trained on Getty's images without their permission and tried to commercialize it via DreamStudio. [0] https://www.technollama.co.uk/high-court-rules-that-getty-v-stability-ai-case-can-proceed https://www.technollama.co.uk/high-court-rules-that-getty-v-... [1] https://www.theverge.com/2023/1/17/23558516/ai-art-copyright-stable-diffusion-getty-images-lawsuit https://www.theverge.com/2023/1/17/23558516/ai-art-copyright...
- HarHarVeryFunny 3y agoThe issue is the output, not the input. If LLMs were just learning from their training set and generating novel output, the same way a human might, then there'd be no problem. The issue is that generative AI - both for images as well as text - in effect memorize sources as well as learn from from, and can end up regenerating training sources verbatim (or with minimal changes in case of images). I don't think any US court is going to accept "yes your honor, we copied this copyright material, but we used a TOOL to do it" as a way to avoid copyright.
- badcppdev 3y agoI don't understand your point because humans generate novel output AND memorise and regurgitate source material. I believe human artists consciously adjust their output to avoid copying previous artists too closely. And sometimes they choose to copy very closely or exactly. Obviously the same feature can be implemented as an option on generative AI systems.
- HarHarVeryFunny 3y agoMy point is that Japan can claim whatever they want about copyright not applying to LLM inputs (training data), but it makes no difference. Copyright infringement will be judged on what they output. Of course, many outputs will be novel enough to avoid copyright claims, but not all, same as the output of a human. No doubt generative AI systems could be built to self-police and not emit any tentative outputs that are too close to training samples, but that's certainly not the way they are today, and it's not clear from the article that started this topic that this issue is addressed in any way by the Japanese law, which is just about training data.