5 ms·
Scale matters. > But suppose a human [...] A human doesn't ingest half of the web and simultaneously deal with millions of people. We've be been through th
by makeitdouble 1y ago
Scale matters.
> But suppose a human [...]
A human doesn't ingest half of the web and simultaneously deal with millions of people.
We've be been through this times and times again. Justice didn't go after humans copying books by hand, it went after reprints of existing copyrighted material.
Music industry didn't go after people singing tunes in their kitchen, but after wide distribution networks.
Removing scale from the discussion leads to absurd conclusions.
- spwa4 1y ago> Justice didn't go after humans copying books by hand, it went after reprints of existing copyrighted material. Not sure who "justice" is, but copyright owners most definitely did go after individuals copying even small excerpts of copyrighted material, in fact that is currently still going on in libraries: https://www.theguardian.com/commentisfree/2023/oct/09/us-library-system-attack-digital-licensing https://www.theguardian.com/commentisfree/2023/oct/09/us-lib... (yes, they explicitly made taking excerpts impossible and fought against any attempt to change that) > Music industry didn't go after people singing tunes in their kitchen Again, yes they did. I'm not aware of kitchen incidents, but this happened: https://stanforddaily.com/2020/08/04/warner-chappell-music-sues-people-for-singing-happy-birthday-while-washing-their-hands/ https://stanforddaily.com/2020/08/04/warner-chappell-music-s...
- makeitdouble 1y ago> individuals copying even small excerpts of copyrighted material Legal copyright claims (not the YouTube kind) need to justify harm to the original piece. > libraries Setting aside my opinion on the situation of public libraries, I wouldn't call libraries "humans copying books by hand" > this happened: [...] "This article is purely satirical and fictitious. "
- gwd 1y agoI agree that we have a new situation developing; but we're not going to get any clarity unless we see clearly what the new situation is. There are several things you're still conflating: 1. The ability of an entity (human or AI) which can be prompted to produce copyright-infringing material. 2. Actually producing copyright-infringing material. Sure, if OpenAI is actually producing copyright-infringing material at scale, unprompted, then that needs to be addressed. If a common way around NYT's paywall were to copy & paste the first few lines into ChatGPT and then read the rest of the article, then yes, that's a hole that needs to be filled. But that's with the production and dissemination, not the training. Regarding scale, yes, there is a difference here, but it's more subtle than you think. There are probably millions of toddlers who can recite The Gruffalo nearly verbatim. However, each of those toddlers were trained individually. Similarly, there are probably thousands, maybe tens of thousands of artists who, when prompted, could generate an image that would be similar enough to the presented image to violate copyright. But again, each of those individuals were trained separately. The difference that modern tech companies have is that they can train their systems once, and then duplicate the same training across millions of instances. One potential argument to make here would be to say: Training with this material is fair use; but fair use or not, those weights are now a derivative work. You can use exactly one copy of those weights, but you can't copy those weights millions of times, any more than it would be fair use to distribute one copy of that image to everyone in the company. You need to either train millions of copies, or pay licensing fees. I'm not sure I agree with that argument, but at least it seems to me to bring the actual issues into more clarity.