4 ms·
I am hard put to defend reddit and a lot of the things they do ... but I can attest, from personal experience, that scrapers are a huge problem for anybody tryi
by corbet 2mo ago
I am hard put to defend reddit and a lot of the things they do ... but I can attest, from personal experience, that scrapers are a huge problem for anybody trying to keep a complex site working. Withholding the display of the comments will significantly reduce the database traffic needed to satisfy a request; I can see why they would want to do that for anonymous readers.
- rob 2mo agoAgreed; we have a Linode VPS for some small WordPress websites and it was grinding to a halt after an AI crawler started iterating through every possible date in "The Events Calendar"’s URL structure, sequentially scraping calendar pages all the way back to like 1970. I'm sure that's partially my fault for not putting it behind Cloudflare or anything though.
- account42 2mo agoThe response to that should be to restrict the events calendar and similar dynamic functionality, not to restrict the 99% of users who are accessing what should be 100% staticly cached HTML.
- rob 2mo agoGood idea; I'll make sure we have all random ?date=1970-01-01, ?date=1970-01-02, all the way to ?date=2026-07-22 URLs cached, even the ones that are 404s but the crawlers still try anyways. Thank you!
- dormento 2mo agoOf course, the _actual_ response to that should be to make the AI companies somehow liable for the massive disruption they caused (and still cause). But alas, we can't have nice things.
- inigyou 2mo agoThere's no evidence this massive DDoS attack is related to AI.
- inigyou 2mo agoWell Reddit would have a lot less scraper traffic if it hadn't shut down the feed API 3 years ago in an attempt to keep its data closed and only distributed to Google (which paid money for it). Self-inflicted. No sympathy.
- corbet 2mo agoI doubt it. The scraper botnets aren't going to bother with niceties like APIs, they just come in over HTTPS and grab anything they can.
- inigyou 2mo agoThere was a whole mini-industry around scraping specifically Reddit using the API feed because it was so open. AI companies first trained on Pushshift, which is like Common Crawl for Reddit, which is why Reddit shut down Pushshift (with legal threats IIRC) at the same time.
- Hizonner 2mo agoIf there were a well-known URL that retrieved all content added to a site after a given date, and a meaningful number of sites implemented it, I bet the biggest scrapers would use it. Except, of course, that then sites would start lying and sending abridged, poisoned, or nonexistent data. Nice things cannot be had, no matter how many levels you go down.
- inigyou 2mo agoReddit used to have something like that. And what you said happened. They shut it down, so that they could ask AI companies for payment.
- uhhhhwhaaaa 2mo agoI personally don't like it, but hard to blame them. Outside some Stoic altruists, this is their advantage and I'd do the same.
- NitpickLawyer 2mo agoI don't understand why sites like reddit need to do any work on (unauthd) GET. They are not instant messaging platforms, they are forums. You can live with a ~1-2-5 minute delay on forums. So POSTs (from logged in users) append to a queue, and a worker does work at regular intervals, creates the HTML and that gets sent to a cache + distributed to a CDN if needed. There's no need to do work on unauthd GETs for a forum.
- cole-k 2mo agoI know I'm not arguing with you personally here since it seems we both have a bone to pick with Reddit, but I don't understand why they can't use Anubis or whatever captcha thing I may or may not have seen on New Reddit. It irks me that they've decided they will keep hosting it (for now), but they're just going to close the gate. And dang it, I swear to god they put in a 2s rate-limit on the page, which it seems they have left in after the transition (perhaps because it's being rolled out still, although they could be so kind as to not rate-limit me now that I've logged in). Was that all they tried??
- Terr_ 2mo ago> I don't understand why they can't use Anubis or whatever captcha thing I may or may not have seen on New Reddit. This is just the first step in killing old-Reddit for logged-in users as well. It's too easy to use, they make more money with the enshittified new version.
- Terr_ 2mo agoExcept comments are really the only reason to visit! Sharing cool links was already old with Digg etc. Worst still, enshittified "new" Reddit will serve a page that claims to have your search terms in it, except the alleged result is buried behind an undetermined path of a tree of incremental "load more comments" button-links... and at some depth those silently switch to causing a whole page-transition to a subtree, and if you navigate back you'll find everything is collapsed/unloaded again and you can't tell which branch you were on before the interruption. With classic Reddit, you just... `ctrl-f` for the term.