6 ms·
> or cherry-picked their examples from many attempts even if that were true, is it really a defence? > It said AI tools have to incorporate copyrighted works
by Baldbvrhunter 3y ago
> or cherry-picked their examples from many attempts
even if that were true, is it really a defence?
> It said AI tools have to incorporate copyrighted works to “represent the full diversity and breadth of human intelligence and experience.”
that sounds like an appeal "our product isn't much use if we can't violate copyright"
- unsupp0rted 3y agoThe government gets to violate intellectual property in the interest of national security whenever it deems fit. I take building the first AGI to be in the same category.
- Baldbvrhunter 3y ago> I take building the first AGI to be in the same category. That's quite the leap. Some don't even think AGI is possible. And some of those that do, don't think LLMs are on the path. Even if we assume it is, there is a significant amount of non-copyrighted text available to train with. The difference being ChatGPT needs text that provides value in the ChatGPT product for the general audience.
- Slartie 3y agoIf you don't happen to be "the government", it's not exactly in your power to decide that.
- unsupp0rted 3y agoIt’s in the government’s power to decide that AGI is worth carving out gaping exceptions for. Even for private companies. If we were to get AGI in a few years, the ends would absolutely justify the means.
- Baldbvrhunter 3y agoAll that needs is for someone to demonstrate the viability of AGI in those circumstances. Is all that holds back AGI the volume of data? If so, how much data is needed?
- unsupp0rted 3y agoAll that holds back AGI is probably not the volume of data. We're still missing key discoveries. But giving an LLM loads of data might turn out to have been a necessary condition on the road to developing AGI.
- Arnt 3y agoYes, that's a defense. Twice over. I bought an MP3 album from Amazon last weekend. One of the many things I got from that purchase was the ability to copy that album, which would be a copyright violation. That doesn't make the purchase unjustifiable, immoral or illegal — my actual use for the album justifies the purchase. The possible copyright violation is irrelevant. People will try to trick you with statements that mention something bad and omit everything good. Don't let them. Think about what's omitted. Does chatgpt get anything good, legal, useful from reading NYT? I'd say it does. For example, it gets the knowledge necessary to explain things in three paragraphs, partly based on NYT articles. And partly based on Wikipedia, which in turn is based on the NYT. OpenAI is saying that training to providing a three-paragraph summary of recent events is fair use of newspapers, and that such training is not realistically possible without copyrighted materials. It's saying that if you make copyright violations impossible instead of difficult, then you can't use the articles fairly either. Sounds persuasive to me. There's a second aspect, less important IMO: de minimis non curat lex. "The law does not concern itself with trifles" basically. If OpenAI made it really difficult to make GTP do a certain thing, if you have to try many times and it's not even clear whether each attempt succeeded, then the possibility of doing that thing isn't a matter of law, says that principle.
- MattDaEskimo 3y agoYour comparison to the Amazon mp3s doesn't match at all the current, unprecedented situation. The obvious issue is that copyrighted material was used without permission, and an opt-out feature was introduced without any way to remove already used training data. If the copyrighted data that OpenAI is copying is causing financial harm to the rightful owner there is certainely grounds for copyright infringement. This is not fair use. It is not fair to try and draw metaphors.
- Arnt 3y agoWe use copyrighted materials all the time without permission. You and I both read the Verge article without the Verge's permission. Reading is the intended and most common use of the Verge's articles, and neither of us asked for permission. I didn't print that one but I often do print, always without asking anyone's permission. Copyright has that name because copying is exceptionally protected; general use is not. You can argue that training is a kind of copying, since it involves copying of things from RAM to RAM, etc. I find that difficult, since we've established that e.g. this browser's copying of web page contents from RAM to RAM isn't. If you don't argue that training is copying, then you can argue that since training is a necessary prelude to copying, it should be treated like copying legally. I disagree, because various kinds of fair use also has the training as a necessary prelude (and, uh, the purchase I mentioned could also be a necessary prelude to copying, if my goal was to copy the album).
- justanotherjoe 3y agoit's an appeal to the common good. Which is persuasive and always been how i see it. I think everyone here is motivated by greed but there is a huge common good from chatgpt.being able to read news articles. Off the top of my head: fake news detection.
- dangus 3y agoIt definitely isn’t a good defense. This would be like if you were able to say some magic words to goad Google Search into giving you a copy of Avatar: The Way of Water. Even worse, the movie would be a file hosted and distributed by Google directly.
- gooob 3y agothey shouldn't of ever made it into a "product". it wasn't ready. this is still a new technology and a new tool in a unfinished early stage in its development. it is still in testing phase. companies shouldn't be trying to make money off of it yet.