5 ms·
I think it's very clear that the recent LLM boom is directly responsible for Twitter, Reddit, and others quickly moving to restricted APIs with exorbitant prici
by 58x14 3y ago
I think it's very clear that the recent LLM boom is directly responsible for Twitter, Reddit, and others quickly moving to restricted APIs with exorbitant pricing structures. I don't think these orgs really care much about third-party clients other than a nuisance consuming some fraction of their userbase.
Enterprise deals between these user generated content platforms and LLM platforms may well involve many billions of API requests, and the pricing is likely an order of magnitude less expensive per call due to the volume. The result is a cost-per-call that is cost-prohibitive at smaller scales, and undoubtedly the UGC platform operators are aware that they're pricing out third-party applications like Apollo and Pushshift. These operators need high baseline pricing so they can discount in negotiation with LLM clients.
Or, perhaps, it's the opposite: for instance, Reddit could be developing its own first-party language model, and any other model with access to semi-realtime data is a potentially existential competitor. The best strategic route is to make it economically infeasible for some hypothetical competitor to arise, while still generating revenue from clients willing to pay these much higher rates.
Ultimately, this seems to be playing out as the endgame of the open internet v. corporate consolidation, and while it's unclear who's winning, I think it's pretty obvious that most of us are losing.
- amelius 3y agoCan't they pull the data from archive.org?
- notacoward 3y agoThat would be worse.
- KuiN 3y agoArchive.org was knocked offline the other day due to some AI startup scraping it to death. It’s not a good thing.
- moneywoes 3y agoSource, they don’t rate limit
- Kon-Peki 3y agohttps://news.ycombinator.com/item?id=36110527 https://news.ycombinator.com/item?id=36110527
- piperswe 3y agoTrue - and their lack of rate limiting ended up letting someone overwhelm their servers, knocking them offline.
- edgyquant 3y agoThey put out a blog asking people not to scrape afterwards. A simple google will be much fast than asking for sources.
- SllX 3y agoArchive.org is a non-profit without the capacity to serve that many requests. An excellent resource for people to use carefully, but not a treasure trove for bots to scrape down to the last bit.
- notpushkin 3y agoWould be cool if they introduce some reasonably priced access for mass scrapers. Should make some nice income in addition to donations, and a valuable service to community.
- Nextgrid 3y agoLLMs have nothing to do with it. Someone skilled enough and rich enough to develop and train an LLM is absolutely capable of reverse-engineering your private API or scraping your web UI and defeating whatever protections you have.
- throw_nbvc1234 3y agoAnd open yourself to potential lawsuits. You can fork any public repo in github too, don't need any fancy resverse-engineering or web scraping. But if you use the content illegally then what's the point.
- dvngnt_ 3y agoi think you can only scrape public information, so if everything is behind a login screen then that might cause issues
- moneywoes 3y agoManaging all those LTE proxies is far from cheap
- numpad0 3y agoI heard researchers on public funding can’t violate ToS without invalidating their current and future employment, and therefore cannot engage in social media researches without free API…
- appleaday1 3y agodont tell this to youtube
- lost_tourist 3y agolol you can't get in trouble for datascraping or figuring out ways around their anti scraping measures. Good luck enforcing any user agreements the bot has to click through. If they don't want it scraped then they have to not put it on a public facing webpage.
- quartz 3y agoYes it's this. This has nothing to do with 3rd party app operation and everything to do with generally closing the gate to the data garden. The value of reddit's content to non-reddit entities is rapidly increasing as its monetizable use shifts from a set of signals on which to build first-party ad targeting (which they never really figured out) to generally useful llm training data.
- eru 3y agoIf you want training data for an LLM and are actively talking to some data providers, you'd just ask for a dump, instead of making a billion small requests. (You'd make the billion small requests, if you are doing this on the sly.)
- sahila 3y agoRight that'd be the case now but previously you could just make a billion small requests for free.
- eru 3y agoOr at least you could try. But that still makes the original commenters argument moot: > Enterprise deals between these user generated content platforms and LLM platforms may well involve many billions of API requests, and the pricing is likely an order of magnitude less expensive per call due to the volume. The result is a cost-per-call that is cost-prohibitive at smaller scales, [...] That speculation is not how things have been or were.
- fluidcruft 3y agoI think most people who wanted large datasets got their data via pushshift. Pushshift was basically a guy who started out doing small things got so frustrated with the API that he eventually grew to maintaining large mirrors of Reddit content on Google cloud that people could access and query. I don't know why anyone doing research would have used reddit's API instead of using pushshift. Pushshift has been shutdown by reddit earlier this year, so probably they are getting hammered by LLM folks trying to get the data now since they killed pushshift without understanding how it fit into the universe. Reddit is completely stupid if they think people are going to pay for "enterprise API" access... pushshift existed because the API was trash and the only real option is to dump the entire dataset into something usable. The reason reddit's data was used so much is because there was an SQL API via pushshift and you could also download archives of the entire dataset at one go.
- 3y ago
- dageshi 3y agoThis is very obviously what's going on. The web is in the process of rapidly filling up with AI regurgitated garbage, eventually there's going to be a handful of sites with real usable content on them left, reddit being one of the biggest.
- xtracto 3y ago>The web is in the process of rapidly filling up with AI regurgitated garbage This is already the case. See the oceans of crap SEO optimized "food recipe sites". It's unbearable. So sad that, BBC back in 199ps and 2000s, there were so many random sites to visit with interesting things. Search engines were of actual use. Now, it's basically facebook, reddit, pinterest, instagram, stackoverflow , and a couple of counted others, depending on what you like. And EVERYTHING is monetized. The WWW of today is terrible. Now
- hackernewds 3y agoah that explains why Twitter led the pack making their APIs insanely expensive. there is value in the data and the LLM companies will be willing to fork it. a whole new business model and monetization of mass data not predicated on ads or user privacy. what could go wrong?
- Macha 3y agoI don't think the LLM boom caused Twitter's first API lockdown in 2012, nor do I really think it's anything to do with the more recent final nail which seems much more in tune with Elon's twitter trying to increase ARPU/engagement while also dealing with a 90% reduction in headcount.