7 ms·
I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corp
by eggbrain 3y ago
I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took:
- The entire corpus of data the community had curated over the last XX years
- The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data
- Potentially large amounts of traffic that would normally come to their sites via Google (e.g. site:reddit.com), that is now available instantly (and customized) via ChatGPT
Despite Reddit's probably closer connections to OpenAI than other startups through Y-Combinator and Sam Altman, I wonder how keen they are to actually work with a company that potentially destroyed a ton of their value, right before they were ready to IPO.
- skilled 3y agoI agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenever you train the next model. [0]: https://old.reddit.com/r/reddit/comments/145bram/addressing_the_community_about_changes_to_our_api/jnk9izp/?context=3 https://old.reddit.com/r/reddit/comments/145bram/addressing_...
- ltadeut 3y ago> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?
- EMIRELADERO 3y agoIt is OpenAI's and Microsoft's position that obtaining, using for training, and displaying data via AI is fair use.
- asylteltine 3y agoI really really really hope the copilot and other suits are successful. The idea that you can literally steal content in the name of “””AI””” and profit off it is just insane! How is it not like copyright infringement? The Warhol case is just one step behind training data. It’s basically the same idea.
- skilled 3y agoInteresting question but sadly I am in no position to answer it. I think there are probably issues to address with scraping it blindly: - Can Reddit imprint its data somehow? A watermark? - Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization? If OpenAI can't work around this, I'm not sure they would be willing to cross any lines in terms of copyright, they've already done it with ChatGPT and I am guessing rules are only going to get stricter on this topic.
- comfypotato 3y agoI think the bottom line is that Microsoft’s (and thus other for-profit AI initiatives) stance is that any and all data is fair game regardless of license or authorization. This results, in their opinion, from the fact that the AI alters the data, changes the output, and is otherwise “inspired” by the data in the same way an artist might be inspired by another without copyright infringement.
- nologic01 3y agoThis sounds very dodgy. Will somebody be checking the degree of such "data alteration" and verify that the "AI" is actually inspired rather than copying? To me this feels like its opening up the door for the elimination of copyright as any algorithmic layer interjected between scrapped data and end users could claim to be "inspired".
- comfypotato 3y agoWelcome to the discussion lol. People have already provided examples of copilot producing niche code verbatim thereby proving their intuition incorrect. It’s a whole mess that will take years to be cleaned up by new legal conventions.
- Jevon23 3y agoIf my compiler was “inspired” by leaked Windows source code and altered it into a new form then I think their opinions on the matter would be very different.
- csdvrx 3y ago> Wouldn't that be a violation of Reddit's terms and liable to some legal procedure? It won't, with the LinkedIn vs HiQ precedent.
- dom96 3y agoHow does Google use Reddit's data in its models? You can access most (all?) Reddit pages without hitting Reddit at all via the "Cached" link in the search results. Does Google have a special agreement with Reddit (and all other sites?) or is it legally "fair use" to reproduce web pages that are available freely online?
- mminer237 3y agoI think that's a different legal question than LLM training, but webpage caching has been found to be fair use based on a number of factors: https://www.pinsentmasons.com/out-law/news/google-cache-does-not-breach-copyright-says-court https://www.pinsentmasons.com/out-law/news/google-cache-does...
- mminer237 3y agoReddit's terms are irrelevant. Unless Reddit requires a login to view its site (which would also prevent Google indexing), anyone can view the data without agreeing to the terms. The only question is copyright, but I find it hard to argue that LLM training is not sufficiently transformative in 99% of cases.
- icebraining 3y agoPlus Reddit doesn't own the copyright to the posts, the users do.
- deleted 3y ago[deleted]
- tedivm 3y agoIn the US at least the courts has made it clear that scrapping is legal. https://techcrunch.com/2022/04/18/web-scraping-legal-court/ https://techcrunch.com/2022/04/18/web-scraping-legal-court/
- cma 3y agoReddit don't own the copyright to it, just a license. That plus public web scraping is legal. Reproducing the data directly might violate the user's copyrights, but through an LLM it is assumed not.
- spacephysics 3y agoYeah I think their real intention is to kill off third party Reddit apps so people are forced to use their own app, with all its tracking and garbage. Reddit on mobile browser is a case study of insane dark patterns Click to sort comments while not logged in? A popup appears asking to log in, with no close button. You have to click out of the box, but that’s not easily apparent View an 18+ subreddit? Let them browse for 30 seconds, then ask they log in Visit the site? Ask to login or use their app. At this point, I’m more motivated to do anything but use their app if they’re this hellbent on getting me to download it.
- le-mark 3y agoThis exactly, it’s unuseable and will only get worse. When “old” stops working it’ll be a sad, sad day. This blackout has made me realize how deeply I’ll miss reddit. The thing is, to me, reddit would never be a viable enterprise without the volunteers moderation. If reddit had to pay for that they’d never ipo. And if they think the moderators will stay and be exploited they’re wrong. Thus if they go through with this they’ll fail as a public company. There is no there where they going.
- simias 3y agoYeah they've been trying to kill old.reddit and third party apps for a long time. The experience degrades more and more as they add new features that are only fully supported (or supported at all) on new.reddit and the official app. This has been going on for years. I'm convinced that this would have happened with or without OpenAI, especially with the mirage of an IPO on the horizon. Controlling the client to show ads and siphon data is just too valuable. Maybe the OpenAI thing pushed them to speed up the process.
- adql 3y agoI think he's bullshitting even on API costs to maintain. They could just put it into reddit premium "want to use API ? Pay few bucks and use app of your choosing". Even $2 would easily cover the cost of lost ads and such. Then a much more expensive tier above for "data ingestors".
- dontupvoteme 3y ago>Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. The irony if lowtax and :10bux: ends up being right all along about how to build a community...
- DSMan195276 3y agoI'd also point out, it's not like the pricing of the API is going to be uniform for everyone anyway - spez has already said/conceded that certain users of the API can continue to use it for free (above the free limits). At that point it's a deliberate decision to make 3rd party apps pay the same price as AI companies, if they wanted to keep 3rd party apps around they could have set a different more reasonable price point for them or went another route (Ex. require 3rd party app users to have Premium).
- deleted 3y ago[deleted]
- hn2017 3y ago[dead]
- dontupvoteme 3y agoWhy are they overlooking their biggest value as an organization is to sell limited truth to AIbros who want the _true_ up/down vote information, rather than the fake number they decide to provide to the plebians?
- keyringlight 3y agoOne thing that's struck me which I assume only reddit has, is the social graph data which holds all the relationships trends on who votes on what/who. There's limits on that based on how pseudoanonymous it is and the ease of making new accounts, but that seems valuable (at least to an outsider), although possibly in a "fighting yesterday's war" way as 'true' social networks like Facebook get value out of posts, which would be different to the way ML training may value it.
- duxup 3y agoI think the problem is that the votes, not necessarily representative or anything “true” either.
- WORMS_EAT_WORMS 3y agoAh, maybe. I see it differently though. reddit has never been about authentic content. - When they started, I believe they boasted creating fake accounts to mimic engagement to help grow community - It is painfully obvious the amount of political, gov sponsored, and corporate astro turfing that “gets through”. Both human and bot comment farms are real and have no doubt been artificially bolstering ideas and content for years As I see it… They have enabled this / looked the other away at this behavior forever in exchange for engagement. The super high tech AI chat bots are probably going to be welcomed for their clever content (hence why CEO could care less about its users anymore). Their real value has always been having a controlling voice and being able to push viral ideas. The data thing is all hype noise IMO. It was never going to be a serious part of their IPO. Big messy gross company.
- quickthrower2 3y agoThere is a decent amount of good stuff on reddit. With a RLHF model trained on shitposting, political grandstanding, porn etc and use that to filter out, you are left with great stuff. r/experienceddevs. I would fine tune on that filtering out negative voted comments and you have a career advisor. Mods did the data cleaning already!
- theknocker 3y ago[dead]
- bbor 3y agoEh I really strongly suspect this “you can use an LLM instead of google, I.e. as a knowledge model” to be a short lived trend. I hope Reddit sees it the same way. It’s kinda like using better bike infrastructure for sick cyberpunk roller derbies; a nice unexpected use case, sure, but its not built for that and sooner-or-later the issues will become all too apparent IMO, the future will be using LLMs with live search results - which of course will probably require a funding model better than the terrible Display Ads setup we have for such content now. So best of luck to Reddit on that front - I hope your ipo fails
- vineyardmike 3y ago> the future will be using LLMs with live search results N of 1, but I vastly prefer the Google Generative Results over chatGPT. Quality may not always be the same, but a “chat” seems like an awkward metaphor for finding info, and of course, chat GPT has no links to content when I’m worried about hallucinations. I’m true google-search fashion, GenSearch has a lot less deep-in-the-weeds technical answers and will push you to simpler results. Eg. If you want to know what chemicals have a similar absorption properties to methane… you’re better off with ChatGPT or traditional search.
- gsatic 3y agoGoogle or Reddit don't matter, if the Info Explosion problem that the Internet produced gets solved some other way. Main reason such sites came into existence was there was too much info on the Internet. These sites where attempts at simplifying the numerous websites and blogs ppl had to manually discover and track. They have meandered around that problem, and got totally distracted by all kinds of other problems (many self created).
- Aerbil313 3y agoIt’s interesting to think about an information-deduplication solution for internet. Maybe now we can use LLMs and embeddings to create an archive of all (factual) information, then build a search on top of that?
- lisasays 3y agoIf the Info Explosion problem that the Internet produced gets solved some other way. Sure it will - asymptotically. But it took 18 years to grow Reddit. Even if a solid alternative emerges in a fraction of that time -- that's still a multi-year gap. Plus we're at a significant knowledge deficit (for us lowly humans, never mind LLMs) if Reddit's archives don't re-emerge sometime soon. (Google search has been braindead for some time now -- so we can leave that source out of the equation).
- mbmjertan 3y agoGiven that sama is a board member of Reddit Inc, and that this is happening after GPT-4 was trained on Reddit data, I wouldn't jump to conclude they're upset at OpenAI. SO had publicly available, no-auth-required data dumps. This makes it difficult for them to know who is using their data. However, this surely isn't the case for Reddit who offered only API endpoints for this content, and I'm guessing you couldn't use .json to get the whole site (rate limits, etc). I wouldn't be inclined to believe that Reddit would miss a new major API user. This is purely speculation disregarding Hanlon's razor, but I'm thinking that the API pricing comes down to killing two birds with one stone. * Sama got to train his LLM on Reddit and some best-of-the-internet content there such as r/bodyweightfitness, informed discussions on niche topics etc for free. The catch-up players face prohibitive pricing. * Third-party apps get killed, bringing the UX to Reddit's control. I think this is more important to Reddit than ad revenue, as they could've simply built an SDK for probably less than this PR nightmare will cost us. * Their new development platform, however, hints at the Reddit app supporting serving "redditor-made apps" which "can be seamlessly reused between communities". The description (and the idea to have apps in your app) weirdly reminds me of WeChat apps, and given that Tencent is a major shareholder in Reddit, I would consider the possibility that apps are something they're pushing. That idea has no chance of success without the UX being completely in Reddit's hands, even then it's questionable how it would work on Reddit. Spez couldn't actually be that detached from Reddit?
- Aerbil313 3y agoOh shit. Never thought of it that way. Sam is playing real 4D chess here. I’d add that he is also guaranteeing (to some extent) the quality of his data source (reddit) by closing API access, because AI bots can’t easily ruin reddit data without API access.
- jedberg 3y ago> and given that Tencent is a major shareholder in Reddit Snoop Dogg owns more of reddit than Tencent. The Tencent investment is very overblown.
- HaZeust 3y ago
- choas 3y ago- suddenly after the ChatGPT success, they realized that they have valuable data - next step is to stop third-party apps that generated these data - then they let the moderators show their power I’m not sure if people care about a CEO being exposed as a liar nowadays, but maybe some former Yahoo managers have another idea on how to destroy more value.
- rcxdude 3y agoIndeed. Their API changes make very little sense if the goal is to capture the value that content they host has for training LLMs, especially because I don't think there's much value that they can capture there (AI researchers argue that using the content is fair use, and they can simply scrape it if they don't have API access: something which the courts have allowed in similar cases)
- LordDragonfang 3y ago>former Yahoo managers Specifically the ones that ended up in charge of tumblr, so they can suggest "ban adult content on a platform famous for its adult content"
- OldManRyan 3y agoPart of reddit's API changes were going to prevent NSFW content from being viewed through the API but they reversed course on that.
- ineedasername 3y ago>a company that potentially destroyed a ton of their value I'm not so sure that's the case, for two reasons: 1) I have not stopped appending "reddit" to my searches or stopped visiting StackOverflow or other Stack* sites. ChatGPT is simply now an additional tool, and while there's some overlap in use cases there's also plenty I can do with *GPT that I couldn't with other sites. 2) To the extent that a training corpus relies on these (or any other) data sources, OpenAI has just made them much much more valuable! It's sort of like a mining company discovering that someone has an extremely valuable use case for decades of the mining company's tailings. This "someone" may have been allowed to haul away some of it without payment to the mining company, but now they know its value, now they can build a very lucrative business selling access to what remains, and what will be generates in the future. That's all highly simplified though. It's of course much more complex than this, much more complex than can be captured in an HN discussion, but we can explore the outlines a bit, and even disagreement will reveal more and more of it.*
- anaganisk 3y agoIsn't Sam Altman on their board? I wouldn't be surprised if he was the one who pushed for these API changes, to restrict others from building a competitor to OpenAI.