5 ms·
The timing is interesting. I’m convinced their API’s are either being abused, or will be imminently abused by LLM training. I wonder how much third party apps
by guidedlight 3y ago
The timing is interesting. I’m convinced their API’s are either being abused, or will be imminently abused by LLM training.
I wonder how much third party apps are being caught up in this other issue.
- skeyo 3y agoMy initial thought on the timing is how much it coincides with the timing of Twitter's API becoming ridiculously expensive ($42k/mo). That maybe Reddit thought they could hop on that bandwagon and make some coin. But LLM also makes a ton of sense.
- nicce 3y agoLLM does not make much sense anymore to be fair. Too late. They should have closed the APIs years ago. Biggest parties have already mined the data, which is enough for models for long time. Unless you want the model to find some specific comment yesterday.
- d11z 3y agoIt makes sense to me, but for LLM spam bots, not training. I’m assuming it’ll only increase as time passes.
- CTDOCodebases 3y agoReddit is good for product recommendations and in that regards recency matters e.g ask ChatGPT what ODE you should get for a Sega Saturn. By allowing LLMs low cost or free access to their users data companies like Reddit are essentially helping companies like OpenAI choke their traffic. Over time as people realise it’s quicker to ask ChatGPT a question than it is to post a question on reddit they will start losing content also. Also think about the difficulty of policing content generated by swarms of LLM bots with API access for PR campaigns.
- nicce 3y agoAt which point you need API access, and the crawled indexes are not enough enough? Is Google also required to start paying for API access for indexing pages and showing them in the search results? I am just wondering, where is the limit, since in that case the model might not be trained anymore and instead it is used for similar purpose than search engine. I guess Bing is already doing this, without Reddit API.
- pbj1968 3y agoFenrir.
- CTDOCodebases 3y agoFor price and ease of install sure I would agree. I just asked Chat GPT this very question and it didn't mention the Fenrir. It mentioned the other contender (MODE) but it failed to differentiate it from all the other offerings i.e it has the ability to connect a SATA SSD and hold the entire Sega Saturn library on the device.
- Frost1x 3y agoI'm not sure this is true in all cases. As far as grabbing enough language data to produce useful language, it's probably enough for awhile until cultural and language shifts happen (slang, word usage distribution, new terms, etc.). The cases LLM will need this data in is compiling together more modern useful human knowledge as our knowledge base grows. Information in existing LLMs could shift. This is sort of the issue even academic textbooks deal with when publishing what is considered foundational knowledge: sometimes we discover something new that makes it either not quite correct or invalid. These are the sort of obvious failures and disconnects that should become apparent if training lags behind. LLM services interested in revenue without plans for continuously updating training data are somewhat betting that not too much will change from most end users perspectives for awhile and for some use cases that might be true but the limits of training data over time for public instances of GPT for example have already hindered some. Much of prompting, from my anecdata, needs to take that into consideration as one of the base constraints (does this model even have up-to-date information it could query and dump something useful from). To some degree those training limits also help expose "hallucinations" or interpolation/extrapolation attempts of LLM models. If I know it doesn't have this information in the training set and test the system against it, I can observe how well it interpolates, extrapolates, and is transparent about when that's happening. For example if I ask existing models about new syntaxes and structures introduced in Java 21, most should return something back like it doesn't exist, it lacks that newer information, or something to that effect. If instead it starts producing code samples that it couldn't possibly have knowledge of, then I know it's passing back garbage. If it's being continually updated at some frequency, I'm no longer so sure and it may actually be providing new useful information.
- throw_nbvc1234 3y agoHow many of those biggest parties do you expect to give/sell the data to new players if it'd even be legal for them to do so. Closing the gates can still provide a revenue stream for Reddit, especially if they do it early enough into the hype cycle.
- plagiarist 3y agoLLM trainers can just scrape the content. It would only slow them down. I wonder if reddit imagines they would pay for the data.
- dageshi 3y agoThere's maybe a handful of sites that people explicitly append to their google search results in order to cut out the SEO spam, reddit is probably the one that covers the most subjects. So I think they are thinking that and I don't think they're wrong.
- tensor 3y agoI think they're wrong. I was building a product search app using the reddit api, but with their API pricing it isn't remotely feasible financially so I shuttered it. They may see $$$$ but I suspect they'll find very few takers. Yes, the data is valuable for product recommendations, but not at the price they are asking. And if they ever block traditional scraping then all those "append reddit to google" searches will be gone too. Google is not going to pay them either.
- saynay 3y agoUnless they are entirely delusional (which is possible), they have to had priced it with a specific few customers in mind. If they think they can get a cut of ChatGPT money going forward, I think they would be entirely willing to sacrifice all 3rd party apps and a decent chunk of their mobile users.
- rightbyte 3y agoWhy would you train with the API? Just render the site to scrape it.
- Ekaros 3y agoAfter killing the old reddit. I don't think scrapping is very effective, at least for the comments. Visibility of comments on web site is pretty horrid.
- _-____-_ 3y agoThe API is still there, and it must still be usable without paying, since the mobile app is using it. So setup mitmproxy on your iPhone and sniff the traffic to figure out how it gets its key, and then replicate that in your scraper.
- moffkalast 3y agoNo need, there are already several projects that have been scraping and downloading posts and comments for archival reasons for years. It's just a one click download.
- Sol- 3y agoIf that's the case, whitelisting the most popular apps might be somewhat of a solution, no? Also I think being transparent about the fears that LLMs take from your site without returning any value would probably be met with understanding from the userbase.
- saynay 3y agoOn the flip side, if they wanted to kill the popular apps, why wouldn't they just block them directly instead of doing so indirectly through crazy API pricing?
- philjohn 3y agoThe issue is, Reddit is ad supported. The popular apps bypass showing ads. The problem goes away if the apps charge for a subscription, which covers their API access.
- graftak 3y agoThis could have been solved by only allowing Reddit premium users to make use of the API (they don’t see ads). in addition, the API cost is about 20 times more lucrative than a user seeing ads, which is an insane upsell. Then there’s the issue that Reddit won’t allow NSFW content over the API, which is just total BS in the picture you drew.
- bloopernova 3y agoHow difficult is it to rate limit API requests? Just Fibonacci the increasing slowdown. And require something like a public key to access the API so you can track requests coming from multiple hosts. Then any rate of access above that of a power user can be charged for appropriately. And make it so that activating API access is something that can't be automated so people can't create thousands of API dummy accounts.
- manmal 3y ago> track requests coming from multiple hosts The problem here is that many users are behind CGNAT, meaning many end users share a single IPv4. Unfortunately the days of counting distinct users by their IP(v4) are over.
- bloopernova 3y agoI think that could be mitigated by using a public key of some variety.
- manmal 3y agoEvery user would need one.
- celestialcheese 3y agoWhile the APIs make it significantly easier to ingest for LLM training, scraping will still work. Unless they put 100% of content behind a login-gate, it's currently legal* to scrape and use for derivative works, as long as you have the money and the chutzpah to deal with lawsuits that may or may not come. - https://blog.ericgoldman.org/archives/2022/12/hello-youve-been-referred-here-because-youre-wrong-about-web-scraping-laws-guest-blog-post-part-2-of-2.htm https://blog.ericgoldman.org/archives/2022/12/hello-youve-be...
- saynay 3y agoThis is my read as well. I don't really understand the argument that Reddit wants to kill 3rd party apps, so makes an unaffordable API. If they want to kill 3rd party apps... why wouldn't they just kill them directly? Why leave any ability to purchase API access unless they have a customer in mind that can afford to buy it at that price?