4 ms·
Biased TL;DR: Reddit (notable for having a high stock value from their "selling data" business [1]), Medium, Quora, and Cloudflare competitor Fastly created a s
by 1gn15 1y ago
Biased TL;DR: Reddit (notable for having a high stock value from their "selling data" business [1]), Medium, Quora, and Cloudflare competitor Fastly created a standard to restrict what the reader can do with the data users created, called Really Simple Licensing (RSL). Basically robots.txt but with more details, notably with details on how much you should pay Reddit/Medium/Quora.
While this likely has no legal weight (except for EU TDM for commercial use, where the law does take into account opt-outs), they are betting on using services like CloudFlare and Fastly to enforce this.
[1] https://www.investors.com/research/the-new-america/reddit-stock-rddt-ai-ipo-google-search/ https://www.investors.com/research/the-new-america/reddit-st...
- isodev 1y agoIn other words, a lightweight form of DRM. Here come the reasons why we shouldn’t all deploy CloudFlare and similar as gatekeepers to the web. Is there even one example of a “tech mega corp” that has grown to control more than 1/5 of its market without this circling back to hurt people in some way? A single example?
- PhantomHour 1y ago> While this likely has no legal weight I wouldn't be quite so sure about that. The AI industry has entirely relied on 'move fast and break things' and 'old fart judges who don't understand the tech' as their legal strategy. The idea that AI training is fair use isn't so obvious, and quite frankly is entirely ridiculous in a world where AI companies pay for the data. If it's not fair use to take reddit's data, it's not fair use to take mine either. On a technological level the difference to prior ML is straightforward: A classical classifier system is simply incapable of emitting any copyrighted work it was trained on. The very architecture of the system guarantees it to produce new information derived from the training data rather than the training data itself. LLMs and similar generative AI do not have that safeguard. To be practically useful they have to be capable of emitting facts from training data, but have no architectural mechanism to separate facts from expressions. For them to be capable of emitting facts they must also be capable of emitting expressions, and thus, copyright violation. Add in how GenAI tends to directly compete with the market of the works used as training data in ways that prior "fair use" systems did not and things become sketchy quickly. Every major AI company knows this, as they have rushed to implement copyright filtering systems once people started pointing out instances of copyrighted expressions being reproduced by AI systems. (There are technical reasons why this isn't a very good solution to curtail copyright infringement by AI) Observe how all the major copyright victories amount to judges dismissing cases on grounds of "Well you don't have an example specific to your work" rather than addressing whether such uses are acceptable as a collective whole.
- janalsncm 1y ago> The very architecture of the system guarantees it to produce new information derived from the training data rather than the training data itself A “classical” classifier can regurgitate its training data as well. It’s just that Reddit never seemed to care about people training e.g. sentiment classifiers on their data before. In fact a “decoder” is simply autoregressive token classification.
- orangecat 1y ago'old fart judges who don't understand the tech' If this intended to refer to Judge Alsup, it is extremely wrong.
- PhantomHour 1y agoIt is not.
- visarga 1y ago> but have no architectural mechanism to separate facts from expressions Sure they do. Every time a bot searches, reads your site and formulates an answer it does not replicate your expression. First of all, it compares across 20.. 100 sources. Second, it only reports what is related to the user query. And third - it uses its own expression. It's more like asking a friend who read those articles and getting an answer. LLMs ability to separate facts from expression is quite well developed, maybe their strongest skill. They can translate, paraphrase, summarize, or reword forever.
- PhantomHour 1y agoThis is a baseless assertion of emergent behaviour. > Every time a bot searches We are talking about LLMs by themselves, not larger systems using them. > LLMs ability to separate facts from expression is quite well developed It is not. Whether you ask an LLM for an excerpt of the bible, or an excerpt of The Lord of the Rings, the LLM does not distinguish. It has no concept of what is, and what is not, under copyright.
- 1y ago
- luckylion 1y agoDoes that have any implications on liability for content? They're no longer just a provider, they are re-licensing and marketing content. Are they losing protection?
- ec109685 1y agoIt’s surprising Reddit doesn’t get pushback for reselling their user’s content. The right thing would be for the end users to receive the compensation Reddit is getting from AI companies.
- lotsofpulp 1y agoIt is not clear what makes that the right thing. For example, I have probably saved a decent amount of time and money searching for solutions on Reddit, so would it have been “right” for me to compensate Reddit?
- ec109685 1y agoYou did via ads, and some of that value should go to the commenters.