6 ms·
I think this is an edge case where intuition breaks down. I'm not sure if my intuition is broke or yours. My intuition is that, if you publish something online
by persolb 3y ago
I think this is an edge case where intuition breaks down. I'm not sure if my intuition is broke or yours.
My intuition is that, if you publish something online and let me read it, I can use that knowledge. I can NOT flat out duplicate it for others, but I can still use the knowledge. Training OpenAI is akin to me using something I read. (Hell, most of what I know I obtained this way.)
To me, the only issue is if the model is reproducing text verbatim en masse.
I think my intuition above is basically where my intuition lands for search engines too. The can use datas found on the web, and can show me snippets, but I'd draw a line at them showing me all/most of an article.
- johnnyanmac 3y agoWe go straight back to the arguments of last decade's scraping legalities. Think of my intuition as an all you can eat buffet. In spirit, you can only eat so much as one person and you may bring home some scraps of food you already touched. In theory (if we ignore fine print), you can very much eat out half the buffet, and then take the other half to go. But this latter "approach" breaks down the model of a buffet. Tragedy of the commons. Everyone is a glutton, the buffet can't keep up with making food, business shuts down, no more food for anyone. The internet here is a buffet. Just because you have access to terabytes, petabytes or more of information without paying doesn't mean you can download all those terabytes of info. Everyone starts doing that and the model of the commons that are ad-supported servers breaks down. Sites start to close up behind more anti-copy protection or literal paywalls to make things not-so-freely accessible. And that ends the free internet as we know it. We are indeed starting to see pieces of this happen in real time. >To me, the only issue is if the model is reproducing text verbatim en masse. That's the other issue as well that was discovered with that exploit. You can reguritate the source with enough prompting. This seems to imply that some server out there is in fact just storing hoards of source material on itself. If that server is commercialized and has TBs of copyright material, that seems to go past the fair use argument of "limited use".
- Ukv 3y ago> Sites start to close up behind more anti-copy protection or literal paywalls to make things not-so-freely accessible. And that ends the free internet as we know it. We are indeed starting to see pieces of this happen in real time Sites have been pushing towards "to continue reading, download our app" or "you've reached your free article limit, subscribe for just $X" for a long time, because it's profitable. I don't think there's sense in retroactively placing the blame for this behavior on crawlers - it's a tactic designed for humans, to the extent that it's often implemented in ways entirely ineffective against bots such as putting an element in front of the content, using cookies to track viewed article count, or even intentionally letting bot user-agents through the paywall. Sites are capable of restricting GPTBot/CCBot/etc., or putting a rate limit on crawling, without impacting human readers. > You can reguritate the source with enough prompting. This seems to imply that some server out there is in fact just storing hoards of source material on itself. If that server is commercialized and has TBs of copyright material, that seems to go past the fair use argument of "limited use". The model isn't retrieving content from some storage server during inference if that's what you're thinking, but consider Google Books which does actually do that - storing a huge number of in-copyright books internally on Google's servers and retrieving snippets when searched. As found in Authors Guild v. Google, this is Fair Use - in particular: > > What matters in such cases is not so much "the amount and substantiality of the portion used" in making a copy, but rather the amount and substantiality of what is thereby made accessible to a public for which it may serve as a competing substitute.
- johnnyanmac 3y ago>Sites have been pushing towards "to continue reading, download our app" or "you've reached your free article limit, subscribe for just $X" for a long time, because it's profitable. I don't think there's sense in retroactively placing the blame for this behavior on crawlers And websites have been getting paid for a long time by selling user data (which again is only useful for human data, not robot data). I don't think it's a coincidence that the paywalling of more and more sites came as GDPR and other data privacy regulations started to pop up. Crawlers are one of those smaller symptoms that gets knocked out along the way as well. Can't make money one way, they switch to another (even before taking the pandemic and interests rates into account). >Sites are capable of restricting GPTBot/CCBot/etc., or putting a rate limit on crawling, without impacting human readers. all those captchas I'm being hit with these days tells me otherwise. Maybe it is possible, but perhaps too expensive. Or the capthas themselves are great for training data. I don't know the real reason, but the internet has definitely become most hostile towards my human eyes in an attempt to curb bots. And the worst part is that it doesn't even entirely block out the bots. >The model isn't retrieving content from some storage server during inference if that's what you're thinking I don't admit to know the full details, but some exploits discovered lately seem to suggest that this isn't just "reading a page" and then moving on: https://arxiv.org/pdf/2301.13188.pdf https://arxiv.org/pdf/2301.13188.pdf (PDF warning). Apparently somewhere in the memory there is those raw assets being stored for "reference". Again, I don't know enough to debate on it, but I've seen enough pieces to be scrutinous of how the process currently works. Scraping is a gray area, but storing copyright data in your commercial product is a much harder issue to argue. Youtube screws its creators over backwards to prevent litigation over a very similar concept.
- Terr_ 3y ago> if you publish something online and let me read it, I can use that knowledge I think that's a subtly different angle since it doesn't make any appeal to scale/cost. Sort of like the difference between: 1: "Everybody should be free to use Article X in manner Y without negotiation or payment, as a matter of principle." 2: "When I choose to use millions of Article X'es in manner Y, it just becomes expensive and impractical to negotiate/pay, therefore my situation is qualitatively different and exempt."