5 ms·
Wikipedia is struggling with voracious AI bot crawlers
- graemep 1y agoWikipedia provides dumps. Probably cheaper and easier than crawling it. Given the size of Wikipedia it would be well worth a little extra code. it also avoids the risk of getting blocked, and is more reliable. It suggest to me that people running AI crawlers are throwing resources at the problem with little thought.
- voidUpdate 1y agoMaybe they just vibe-coded the crawlers and that's why they don't work very well or know the best way to do it
- ldng 1y agoMaybe they should just ... not "vibe-code" at all then ?
- voidUpdate 1y agoSounds great to me
- milesrout 1y agoWe shouldn't use that term. "Vibe coding". Nope. Brainless coding. That is what it is. It's what beginners do: programming without understanding what they--or their programs--are doing. The additional use of a computer program that brainlessly generates brainless code to complement their own brainless code doesn't mean we should call what they are doing by a new name.
- wslh 1y agoWouldn't downloading the publicly available Wikipedia database (e.g. via Torrent [1]) be enough for AI training purposes? I get that this doesn't actually stop AI bots, but captchas and other restrictions would undermine the open nature of Wikipedia. [1] https://en.wikipedia.org/wiki/Wikipedia:Database_download https://en.wikipedia.org/wiki/Wikipedia:Database_download
- deleted 1y ago[deleted]
- diggan 1y agoThis has to be one of strangest targets to crawl, since they themselves make database dumps available for download (https://en.wikipedia.org/wiki/Wikipedia:Database_download https://en.wikipedia.org/wiki/Wikipedia:Database_download) and if that wasn't enough, there are 3rd party dumps as well (https://library.kiwix.org/#lang=eng&category=wikipedia https://library.kiwix.org/#lang=eng&category=wikipedia) that you could use if the official ones aren't good enough for some reason. Why would you crawl the web interface when the data is so readily available in a even better format?
- wslh 1y agoI think those crawlers are just very generic: they basically operate like wget scripts, without much logic for avoiding sites that already offer clean data dumps.
- ldng 1y agoThat is not an excuse. Wikipedia isn't just any site.
- wslh 1y agoNot an excuse, a plausible explanation of what's actually happening.
- franktankbank 1y agoAlso plausibly they are trying to kill the site via soft ddos. Then they can sell a service based on all the data they scraped + unauditable censoring.
- marginalia_nu 1y agoI think most of these crawlers just aren't very well implemented. Takes a lot of time and effort to get crawling to work well, very easy to accidentally DoS a website if you don't pay attention.
- 1y ago
- perching_aix 1y agoI thought all of Wikipedia can be downloaded directly if that's the goal? [0] Why scrape? [0] https://en.wikipedia.org/wiki/Wikipedia:Database_download https://en.wikipedia.org/wiki/Wikipedia:Database_download
- netsharc 1y agoSomeone's gotta tell the LLMs that when a prompt-kiddie asks them to build a scraper bot that "I suggest downloading the database instead".
- tiagod 1y agoThis is the first time I'm reading "prompt-kiddie", made me chuckle hard. Jumped straight into my vocabulary :-)
- werdnapk 1y agoTurns out AI isn't smart enough to figure this out yet.
- skydhash 1y agoThe worst thing about that is that wikipedia has dumps of all its data which you can download.
- qwertox 1y agoMaybe the big tech providers should play fair and host the downloadable database for those bots as well as crawlable mirrors.
- microtherion 1y agoNot just Wikipedia. My home server (hosting a number of not particularly noteworthy things, such as my personal gitea instance) has been absolutely hammered in recent months, to the extent of periodically bringing down the server for hours with thrashing. The worst part is that every single sociopathic company in the world seems to have simultaneously unleashed their own fleet of crawlers. Most of the bots downright ignore robots.txt, and some of the crawlers hit the site simultaneously from several IPs. I've been trying to lure the bots into a nepenthes tarpit, which somewhat helps, but ultimately find myself having to firewall entire IP ranges.
- PeterStuer 1y agoWhy not just rate limit? IP range based blocking will likely hit far more legitimate users than you think.
- microtherion 1y agoI know how I can block IPs in my router, but I'm not sure how I can rate limit. And I don't want to rate limit on the server level, because that in itself places load on the server.
- ddtaylor 1y agoIPFS
- lambdaone 1y agoIt's not just Wikipedia - the entire rest of the open-access web is suffering with them. I think the most interesting thing here is that it shows that the companies doing these crawls simply don't care who they hurt, as they actively take measures to prevent their victims from stopping them by using multiple IP addresses, snowshoe crawling, evading fingerprinting, and so on. For Wikipedia, there's a solution served up to them on a plate. But they simply can't be bothered to take it. And this in turn shows the overall moral standards of those companies - it's the wild west out there, where the weak go to the wall, and those inflicting the damage know what they're doing, and just don't care. Sociopaths.
- spiderfarmer 1y agoTruth. I have a platform with over a million photos. It costs me a lot of bandwidth.
- aucisson_masque 1y agoPeople got to make bots pay. That's the only way to get rid of this world wide DDOSing backed up by multi billions companies. There are captcha to block bots or at least make them pay money to solve them, some people in Linux community also made tools to combat that, i think something that use a little cpu energy. And in the same time, you offer an api, less expensive than the cost to crawl it, and everyone win. Multi billions companies get their sweet sweet data, Wikipedia gets money to enhance their infrastructure or whatever, users benefits from Wikipedia quality engagement.
- guerrilla 1y agoThis is an interesting model in general: free for humans, pay for automation. How do you enforce that though? Captchas sounds like a waste.
- jerf 1y agoAny plan that starts with "Step one: Apply the tool that almost perfectly distinguishes human traffic from non-human traffic" is doomed to failure. That's whatever the engineering equivalent of "begging the question" is, where the solution to the problem is that we assume that we have the solution to the problem.
- zokier 1y agoIdentity verification is not that far fetched these days. For europeans you got eIDAS and related tech, some other places have similar stuff, for rest of world you can do video based id checks. There are plenty of providers that handle this, it's pretty commonplace stuff.
- greenavocado 1y agoCareless crawler scum will put an end to the open Internet
- m101 1y agoWikipedia spends 1% of its budget on hosting fees. It can spend a bit more given the rest of their corruptions.
- limaoscarjuliet 1y agoBTW, I do not know why you are getting downvoted, this is a real concern that someone needs to tackle one day.
- briandear 1y agoIt’s getting downvoted because the parent comment aligns with what Elon said about Wikipedia; so it’s a knee jerk reaction. Though the sentiment is factual. Previous discussion: (2022) https://news.ycombinator.com/item?id=32840097 https://news.ycombinator.com/item?id=32840097
- insane_dreamer 1y ago> sentiment is factual the sentiment might exist, but that doesn't mean it's based on facts that discussion you linked to can be broken down into: - people upset because they thought Wikipedia was almost bankrupt and it turns out its not (though Wikipedia never claimed to be in its fund-raising) - people upset because they see too many requests for donations - people upset because Wikimedia execs are getting "high" salaries (though they are much _much_ lower than at private co's) - people upset because they think Wikipedia spreads "left-wing ideologies" None of this has anything to do with "corruption".
- m101 1y agoI think there are a lot of blinkered folks who don't see the corruption with Wikipedia. What wiki are experiencing with bots is probably ubiquitous with every website yet the violins come out for them, one of the most corrupt institutions around.
- jraph 1y agoReal concern or not, this is not related to the discussion at hand, which is AI crawlers hammering Wikipedia, which is related to AI crawlers hammering everything these days. Here's the concern at hand. I would like to read on Wikipedia corruption with quality sources (in a separate HN post, which would probably be successful), but that's not quite on-topic here. Not only it's off-topic and borderline whataboutism, it's also not sourced, so the comment doesn't actually help someone who isn't in the knows. Thus, as is, it's not much interesting and kinda useless. These reasons are probably why it has been downvoted: off topic, not helping, not well researched.
- schneems 1y agoI was on a panel with the President of Wikimedia LLC at SXSW and this was brought up. There's audio attached https://schedule.sxsw.com/2025/events/PP153044 https://schedule.sxsw.com/2025/events/PP153044. I also like Anna's (Creative Commons) framing of the problem being money + attribution + reciprocity.
- chuckadams 1y agoWe need to start cutting off whole ASNs of ISPs that host such crawlers and distribute a spamhaus-style block list to that effect. WP should throttle them to serve like one page per minute.
- amazingamazing 1y agowhat's the best way to stop the bots? cloudflare?
- franktankbank 1y agomake pain for the people they serve.
- briandear 1y agoWhy should we stop the bots? Wikipedia supposedly wants the world to have this free information, a bit is just another way of supporting that goal.
- coldpie 1y agoThe answer to your question is in the article.
- richiebful1 1y agoBecause there's a better way for bots to get the data via the wikipedia database dump. Sending some large zip archives is a lot cheaper than individually serving every page on Wikipedia.
- jraph 1y agoNot at all costs though, including disrupting access to said free information. And this free information is not free from rights to respect neither, it's under CC-BY-SA, which requires attribution and sharing under the same conditions, the kind of "subtleties" and "details" with which AI companies have been wiping their big arses.
- gherard5555 1y agomy guess is the gnome anime girl anti bot captcha
- szszrk 1y agomandatory link: https://github.com/TecharoHQ/anubis https://github.com/TecharoHQ/anubis It's an interesting project, I wish there would be better ways to do that, but I guess we are on war with crawlers for a while already.
- fareesh 1y ago"async_visit_link for link in links omg it works"
- dickfor 1y ago[dead]
- delichon 1y agoWe're having the same trouble for a few hundred sites that we manage. It is no problem for crawlers that obey robots.txt since we ask for one visit per 10 seconds, which is manageable. The problem seems to be mostly the greedy bots that request as fast as we can reply. So my current plan is to set rate limiting for everyone, bots or not. But doing stats on the logs, it isn't easy to figure out a limit that won't bounce legit human visitors. The bigger problem is that the LLMs are so good that their users no longer feel the need to visit these sites directly. It looks like the business model of most of our clients is becoming obsolete. My paycheck is downstream of that, and I don't see a fix for it.
- MadVikingGod 1y agoI wonder if there is a WAF that has an exponential backoff and constant decay for delay. Something like start a 10us and decay 1us/s.
- jerven 1y agoWorking for an open-data project, I am starting to believe that the AI companies are basically criminal enterprises. If I did this kind of thing to them they would call the cops and say I am a criminal for breaking TOS and doing a DDOS, therefore they are likely to be criminal organizations and their CEOs should be in Alcatraz.
- PeterStuer 1y agoThe weird thing is their own data does not reflect this at all. The number of articles accessed by users, spiders and bots alike has not moved significantly over the last few years. Why these strange wordings like "65 percent of the resource-consuming traffic"? Is there non-resource consuming traffic? Is this just another fundraising marketing drive? Wikimedia has been know to be less than truthful wrt their funding needs and spent. https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal|bar|2020-02-01~2025-05-01|(access)~desktop*mobile-app*mobile-web+agent~user*spider*automated|monthly https://stats.wikimedia.org/#/all-projects/reading/total-pag...
- zokier 1y agomultimedia content vs articles. It's easy to see how bad scraping of videos and images pushes bandwidth up more than just scraping articles. The resource consuming traffic is clearly explained in the linked post: > This means these types of requests are more likely to get forwarded to the core datacenter, which makes it much more expensive in terms of consumption of our resources. I.e. difference between cached content at cdn edge vs hits to core services.
- diggan 1y agoThe graph you linked seems to be about article viewing ("page views", like a GET request to https://en.wikipedia.org/wiki/Democracy https://en.wikipedia.org/wiki/Democracy for example), while the article mentions multimedia content, so fetching the actual bytes of https://en.wikipedia.org/wiki/Democracy#/media/File:Economist_Intelligence_Unit_Democracy_Index_2024.svg https://en.wikipedia.org/wiki/Democracy#/media/File:Economis... for example, which would consume more content than just loading article pages, as far as I understand.
- shreyshnaccount 1y agothey will ddos the open internet to the point where only big tech will be able to afford to host even the most basic websites? is that the endgame?
- insane_dreamer 1y ago> "expansion happened largely without sufficient attribution, which is key to drive new users to participate in the movement." these multi$B corps continue to leech off of everyone's labors, and no one seems able to stop them; at what level can entities take action? the courts? legislation? we've basically handed over the Internet to a cabal of Big Tech
- deleted 1y ago[deleted]
- laz 1y ago10 years ago at Facebook we had a systems design interview question called "botnet crawl" where the set up that I'd give would be: I'm an entrepreneur who is going to get rich selling printed copies of Wikipedia. I'll pay you to fetch the content for me to print. You get 1000 compromised machines to use. Crawl Wikipedia and give me the data. Go. Some candidates would (rightfully) point out that the entirety is available as an archive, so for "interviewing purposes" we'd have to ignore that fact. If it went well, you would pivot back and forth: OK, you wrote a distributed crawler. Wikipedia hires you you to block it. What do you do? This cat and mouse game goes on indefinitely.
- iJohnDoe 1y agoNot sure why these AI companies need to scrape and crawl. Just seems like a waste when companies like OpenAI have already done this. Obviously, OpenAI won't share their dataset. It's part of their competitive stance. I don't have a point or solution. However, it seems wasteful for non-experts to be gathering the same data and reinventing the wheel.