23 ms·
It seems like the AI crawlers learned how to solve the Anubis challenges
- jsnell 1y agoThis was beyond predictable. The monetary cost of proof of work is several orders of magnitude too small to deter scraping (let alone higher yield abuse), and passing the challenges requires no technical finesse basically by construction.
- zahlman 1y agoWe need to revive 402 Payment Required, clearly. If we lived in a world where we could easily set up a small trusted online balance for microtransactions that's interoperable with everyone, and where giving others a literal penny for their thoughts could allow for running up a significant bill for abusers, I'd gladly play along.
- logicprog 1y agoMe too. I wouldn't mind Project Xanadu style micro payments for blogs, and it'd both fix the AI scraper issue and the ads issue, and help people fund hosting costs sustainably. I think the issue is taxes and transaction fees would push the prices too high, and it'd price out people with very low income possibly. It'd also create really perverse incentives for even more tight copyright control, since your content appearing even in part on anyone else's website is then directly losing you money, so it'd destroy the public Commons even more, which would be bad. But maybe not, who knows.
- myaccountonhn 1y agoPay to visit would be great, and would force these AI companies to actually pay for their data.
- misswaterfairy 1y agoCloudflare are working on this at the moment. https://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/ There's also an open specification called x402: https://www.x402.org/x402-whitepaper.pdf https://www.x402.org/x402-whitepaper.pdf I would definitely use this to charge US$100,000 per request from any AI company to crawl my site. I would exempt 'public good' crawlers like The Internet Archive though. If AI companies valued at billions of dollars want to slurp up my contribution to the human condition, that's my price - subject to price rises only.
- sumtechguy 1y agoFor someone doing spamming that low level would work well. As their cost is determinatively low to make it work. For someone doing scraping to get data and feeding it to an AI not so much. The AI groups usually have some pretty heavy hitting hardware sitting behind it. They could even break off some hardware that is to be retired and have it munch away on it. To make it non cost effective the calculations would need to be much bigger.
- zahlman 1y agoI actually don't understand who Anubis is supposed to "make sure you're not a bot". It seems to be more of a rate limiter than anything else. It self-describes: > Anubis sits in the background and weighs the risk of incoming requests. If it asks a client to complete a challenge, no user interaction is required. > Anubis uses a proof-of-work challenge to ensure that clients are using a modern browser and are able to calculate SHA-256 checksums. Anubis has a customizable difficulty for this proof-of-work challenge, but defaults to 5 leading zeroes. When I go to Codeberg or any other site using it, I'm never asked to perform any kind of in-browser task. It just has my browser run some JavaScript to do that calculation, or uses a signed JWT to let me have that process cached. Why shouldn't an automated agent be able to deal with that just as easily, by just feeding that JavaScript to its own interpreter?
- deleted 1y ago[deleted]
- joe_the_user 1y agoNear as I can guess, the idea is that the code is optimized for what browsers can do and gpus/servers/crawlers/etc can't do as easily (or relatively as easily, just taking up the whole server for a bit might a big cost). Indeed it seems like only a matter of time before something like that would be broken.
- homebrewer 1y agoI think the only requests it was able to block are plain http requests made over curl or Go's stdlib http client. I see enough of both in httpd logs. Now the cancer has adapted by using a fully featured headless web browser that can complete challenges just like any other client. As other commenters say, it was completely predictable from the start.
- deleted 1y ago[deleted]
- yabones 1y agoMy understanding is that it just increases the "expense" of mass crawling just enough to put it out of reach. If it costs fractional pennies per page scrape with just a python or go bot, it costs nickels and dimes to run a headless chromium instance to do the same thing. The purpose is economical - make it too expensive to scrape the "open web". Whether it achieves that goal is another thing.
- rpcope1 1y agoI'm calling it now, this is the beginning of all of the remaining non-commerical properties on the web either going away, or getting hidden inside of some trusted overlay network. Unless the "AI" race slows down or changes or some other act of god happens, the incentives are aligned that I foresee wide swaths of the net getting flogged to death.
- v5v3 1y agoCould it be a 'correct' continuation of Darwin's survival of the fittest?
- herval 1y agoHasn’t that been the case for a while? I’d imagine the combined traffic to all sites on the web combined doesn’t match a single hour of the traffic to the top 5 social media sites. The web is pretty much dead for a while now, many companies don’t even bother maintaining websites anymore
- weinzierl 1y agoI think the answer for the non-commercial web is to stop worrying. I understand why certain business models have a problem with AI crawlers, but I fail to see why sites like Codeberg have an issue. If the problem is cost for the traffic then this is nothing new and I thought we have learned how to handle that by now.
- MYEUHD 1y agoAbout 3 hours ago the codeberg website was really slow. Services like codeberg that are run on donations can be easily DOS'ed by AI crawlers
- myaccountonhn 1y agoThe issue is the insane amount of traffic from crawlers that DDOS websites. For example: https://drewdevault.com/2025/03/17/2025-03-17-Stop-externalizing-your-costs-on-me.html https://drewdevault.com/2025/03/17/2025-03-17-Stop-externali... > [...] Now it’s LLMs. If you think these crawlers respect robots.txt then you are several assumptions of good faith removed from reality. These bots crawl everything they can find, robots.txt be damned, including expensive endpoints like git blame, every page of every git log, and every commit in every repo, and they do so using random User-Agents that overlap with end-users and come from tens of thousands of IP addresses – mostly residential, in unrelated subnets, each one making no more than one HTTP request over any time period we tried to measure – actively and maliciously adapting and blending in with end-user traffic and avoiding attempts to characterize their behavior or block their traffic. The linux kernel has also been dealing with it AFAIK. Apparently it's not so easy to deal with, because these ai scrapers pull a lot of tricks to anonymize themselves.
- hyghjiyhu 1y agoCrazy thought but what if you made the work required to access the site equal the work required to host site. Host the public part of the database on something like webtorrent. Render website from db locally. You want to ruin expensive queries? Suit yourself. Not easy, but maybe possible?
- Retr0id 1y agoLast time I checked, Anubis used SHA256 for PoW. This is very GPU/ASIC friendly, so there's a big disparity between the amount of compute available in a legit browser vs a datacentre-scale scraping operation. A more memory-hard "mining" algorithm could help.
- jsnell 1y agoA different algorithm would not help. Here's the basic problem: the fully loaded cost of a server CPU core is ~1 cent/hour. The most latency you can afford to inflict on real users is a couple of seconds. That means the cost of passing a challenge the way the users pass it, with a CPU running Javascript, is about 1/1000th of a cent. And then that single proof of work will let them scrape at a minimum hundreds, but more likely thousands, of pages. So a millionth of a cent per page. How much engineering effort is worth spending on optimizing that? Basically none, certainly not enough to offload to GPUs or ASICs.
- Retr0id 1y agoNo matter where the bar is there will always be scrapers willing to jump over it, but if you can raise the bar while holding the user-facing cost constant, that's a win.
- jsnell 1y agoNo, but what I'm saying is that these scrapers are already not using GPUs or ASICs. It just doesn't make any economical sense to do that in the first place. They are running the same Javascript code on the same commodity CPUs and the same Javascript engine as the real users. So switching to an ASIC-resistant algorithm will not raise the bar. It's just going to be another round of the security theater that proof of work was in the first place.
- Retr0id 1y agoThey might not be using GPUs but their servers definitely have finite RAM. Memory-hard PoW reduces the number of concurrent sessions you can maintain per fixed amount of RAM. The more sites get protected by Anubis, the stronger the incentives are for scrapers to actually switch to GPUs etc. It wouldn't take all that much engineering work to hook the webcrypto apis up to a GPU impl (although it would still be fairly inefficient like that). If you're scraping a billion pages then the costs add up.
- logicprog 1y agoI'm not anti-the-tech-behind-AI, but this behavior is just awful, and makes the world worse for everyone. I wish AI companies would instead, I don't know, fund common crawl or something so that they can have a single organization and set of bots collecting all the training data they need and then share it, instead of having a bunch of different AI companies doing duplicated work and resulting in a swath of duplicated requests. Also, I don't understand why they have to make so many requests so often. Why wouldn't like one crawl of each site a day, at a reasonable rate, be enough? It's not like up to the minute info is actually important since LLM training cutoffs are always out of date anyway. I don't get it.
- barbazoo 1y agoGreed. It's never enough money, never enough data, we must have everything all the time and instantly. It's also human nature it seems, looking at how we consume like there's no tomorrow.
- WD-42 1y agoThis is sad, but predictable. At the end of the day if I can follow a link to an Anubis protected site and view it on my phone, the crawlers will be able to as well. I see a lot more private networks in our future, unfortunately.
- jjangkke 1y agoThe private network is only as good as the weakest link which has to offer a reason for people to go through the trouble of accessing either by paying money or other means (acquiring special equipment). And by putting a wall up you end up losing a large portion of the market to those that will now simply arbitrage and fill the space you leave behind. There is simply no way to stop crawlers/scrapers, period, unless you put a meter on it or go offline.
- WD-42 1y agoThe people that are already using Anubis don’t care about “losing a large portion of the market” we just want to work on our FOSS projects without paying unnecessary bills due to out of control crawling.
- electroly 1y agoPresumably they just finally decided they were willing to spend ($) the CPU time to pass the Anubis check. That was always my understanding of Anubis--of course a bot can pass it, it's just going to cost them a bunch of CPU time (and therefore money) to do it.
- zelphirkalt 1y agoI think so too. Maybe the compute cost needs to be upped some more. I am OK with waiting a bit longer when I access the site.
- delusional 1y agoIf I worked at a billion dollar firm, where doing this was actually a profitable endeavor, I'd reimplement the Anubis algorithm in optimized native code and run that. I wouldn't be surprised if you could lower the cost of generating the proof by a couple of orders of magnitude, enough to make it trivial. If you then batch it, or distribute it across your GPU farm, well now it's practically free.
- hollow-moe 1y agoReally looks like the last solution is a legal one, using the DMCA against them using the digital protection or access control circumvention clause or smth.
- jjangkke 1y agoDMCA only applies to hosted content and we've established that LLM aren't hosting copyrighted content as there is significant transformation which you would otherwise need to prove yourself by training and replicating their entire model. There is no legal recourse here, if you don't want AI crawlers accessing your content 1) put it behind a paywall 2) remove from public access
- hollow-moe 1y agoI'm not talking about the output of the LLM here. DMCA is an overreaching law. Here i'm talking about its provisions for access controls and "digital locks", i am not a lawyer but i'm fairly sure you could find some way to categorize Anubis/another software as a digital lock and then sue them on that basis.
- nine_k 1y agoWhy not ask it to directly mine some bitcoin, or do some protein folding? Let's make proof-of-work challenges proof-of-useful-work challenges. The server could even directly serve status 402 with the challenge.
- catsma21 1y agohttps://news.ycombinator.com/item?id=44880393 https://news.ycombinator.com/item?id=44880393
- jjangkke 1y agoDo you want people to mine bitcoin or do protein folding to read your blog or access your web application? More importantly do you want to now compete with those that do not bottleneck and lose your traffic ? This is the paradox, the length you go to protect your content only increases costs for everybody else who isn't an AI crawler.
- nine_k 1y agoPeople, no! Robots which can't pass for people, yes.
- SkiFire13 1y agoNote that the work needs to produce a result that's quickly verifiable by the server.
- Havoc 1y agoReally feels like this needs some sort of unified possibly legal approach to get these fkers to behave. Search era clearly proved it is possible to crawl respectfully - the AI crawlers have just decided not to. They need to be disincentivized from doing this
- dathinab 1y agothe problem in many cases is that even if such a law is made it likely - is hard to enforce - misses bite, i.e. it makes you more money to break it then any penalties but in general yes, a site which indicates they don't want to be crawled by AI bots but still gets crawled should be handled similar to someone with house ban on a shop forcing them self into the shop given how severely messed up some millennia cyber security laws are I wonder if crawlers bypassing Anubis could be interpreted as "circumventing digital access controls/protections" or similar, especially given that its done to make copies of copyrighted material ;=)
- jjangkke 1y agoI really don't get this type of hostility If you put something in public domain people are going to access it unless you put it behind a paywall but you don't want to do it because that would limit access or people wouldn't pay for it to begin with (ex. your blog nobody wants to pay for) There's no law against scraping, and we've already past the CFAA argument
- bargainbin 1y agoIt’s not quite as simple as “putting something in public domain”. The problem is the server costs to keep that thing in the public domain.
- myaccountonhn 1y agoLook at it from a lens of harm rather than legality. The hostility comes from people having to pay thousands in bandwidth costs and having services degraded. These AI companies incur huge costs from their wasteful negligence. It's not reasonable.
- 1y ago
- egypturnash 1y agothanks for making everything that much shittier just so you can steal everyone's data and present it as your own, AI companies!
- amarcheschi 1y agoThe tiniest relief is knowing that models will be distilled and "copied" in smaller models of equal capabilities in ~6/12months, since their output can't be copyrighted and will be used to improve others. Kinda ironic
- black_puppydog 1y agoDear god that invokes the image of some un-human centipede...
- jjangkke 1y agoyou can't steal something that is in public domain and one which you make readily available by publishing it online because there is no provable cost of damage to you by someone scraping and training their models. if you really think what you offer has value, put it in behind a paywall and see how many people will consume it then, probably not a lot.
- xena 1y agoI just found out about this when it came to the front page of Hacker News. I really wish I was given advanced notice. I haven't been able to put as much energy into Anubis as I've wanted because I've been incredibly overwhelmed by life and need to be able to afford to make this my full time job. Support contracts are being roadblocked, and I just wish I had the time and energy to focus on this without having to worry about being the single income for the household.
- veqq 1y agoGood luck, I'm sorry for all of this speculation and people attacking your solution instead of suggesting concrete improvements to help fight the problem.
- xena 1y agoThanks. It means a lot. Today has not been a good day for me. It will be fixed. Things will get better, but this has to rank up there in terms of the worst ways to find out about security issues. It sucks lol.
- ziml77 1y agoI saw you touching grass, so I hope that's at least helping you get through the day <3
- rapnie 1y agoYou are doing a tremendous job, and we are really thankful for the great work you've done. Personal matters come first though imho. Take care <3
- grayhatter 1y agoI'll double down on what veqq said; Thoes that can, do. Those who have no idea where to start complain on internet threads. There will always be bots, they were here before anubis, they'll be there long after you block them again. Take care of yourself first. There's no need to make a bad day worse trying to sprint down a marathon.
- OutOfHere 1y agoThey failed to properly block/throttle the IP subnet as per their admission, and are now blaming others for their failure.
- yogorenapan 1y agoI've seen a lot of traffic from Huawei bypassing Anubis on some of the things I host as well. The funny thing is, I work for Huawei... Asking around, it seems most of it is coming from Huawei Cloud (like AWS) but their artifactory cache also shows a few other captcha bypassing libraries for Arkose/funcaptcha so they're definitely doing it themselves too. Anonymous account for obvious reasons.
- xena 1y agoPlease have someone in the common sense department email me@xeiaso.net. Funding of the Anubis project would go a long way towards mending bridges.
- paddw 1y agoWho exactly is Huawei corporate interested in mending bridges with? Seems like that tie is long severed
- yogorenapan 1y agoI wish I had that kind of power. The European labs get funding from HQ on a per project basis and it takes a lot of effort to convince them to do anything. To even contribute code to open source, we need to fill out a bunch of paperwork, much less fund open source projects that work specifically against some other team's objectives. I'd personally just IP block Huawei's entire ASN since much of it is customer controlled and used for scraping. I know a ton of sites are already doing that since I get IP blocked on my work laptop all the time while doing research with their VPN on
- deleted 1y ago[deleted]
- TZubiri 1y agoIt's PoW, AI crawlers didn't learn shit, their admins just increased their CPU/GPU/ASIC budget
- 1gn15 1y agoDuh. The author of Anubis really should advertise it as a DDoS guard, not an AI guard. Otherwise, xe is just misleading people, while being unnecessarily discriminatory against robots (robotkin) and cyborgs (those using AI agents as an extension of their selves).
- johntash 1y agoIt's not really a DDoS guard though. If someone wants to ddos a server, anubis isn't going to be able to stop the traffic before it gets to the server. It does help from accidental ddos or just rude scrapers that assume everyone has unlimited bandwidth and money.
- bombcar 1y agoWe’ve evolved. The only thing that pays attention to the OSI model is DDoS which can hit you on every layer. And against the application layer attacks Anubis and friends are effective.
- zeropointsh 1y agoHow about using on-chain proof-of-work? It flips the script. If a bot wants access, let it earn it—and let that work be captured, not discarded. Each request becomes compensation to the site itself. The crawler feeds the very system it scrapes. Its computational effort directly funds the site owner's wallet, joining the pool to complete its proof. The cost becomes the contract.
- viraptor 1y agoThe check has to apply to people and bot visitors the same. If you're expecting a blockchain registered spend before the content is visible, basically nobody will visit your website.
- apetresc 1y agoI don’t think OP meant you pay directly, just that you volunteer to do some part of the PoW (of some chain designed for this purpose) on behalf of the site, to its credit. That’s not much of a different ask from Anubis. It just commandeers the compute for some useful purpose.
- delusional 1y agoMaking a proof of work algorithm do some actually useful work is very much an unsolved problem.
- zeropointsh 1y agoExactly, instead of trying to totally prevent the bots/AI scrapers, make them "pay you" in compute to access and scrape. It needs to be solved.
- apetresc 1y agoI don’t necessarily mean “useful” in the sense of scientific research or something. Simply transferring value to the creators of the resource being accessed (a la a blockchain) instead of doing throwaway work would be an improvement.
- chmod775 1y agoFight fire with fire by serving these guys LLM output of made-up news. Wish them good luck noticing that in their dataset.
- johntash 1y agoI think there was some sort of fake webserver that did something like this already. Basically just linked endlessly to more llm-generated pages of nonsense.
- zbentley 1y agoThere are several! Some focus on generating content that can be served to waste crawler time: crates.io/crates/iocaine/2.1.0 Some focus on generating linked pages: https://hackaday.com/2025/01/23/trap-naughty-web-crawlers-in-digestive-juices-with-nepenthes/ https://hackaday.com/2025/01/23/trap-naughty-web-crawlers-in... Some of them play the long game and try to poison models' data: https://codeberg.org/konterfai/konterfai https://codeberg.org/konterfai/konterfai There are lots more as well; those are just a few of the ones that recently made the rounds. I suspect that combining approaches will be a tractable way to waste time: - Anubis-esque systems to defeat or delay easily-deterred or cut-rate crawlers, - CloudFlare or similar for more invasive-to-real-humans crawler deterrence (perhaps only served to a fraction of traffic or traffic that crosses a suspicion threshold?), - Junk content rings like Nepenthes as honeypots or "A/B tests" for whether a particular traffic type is an AI or not (if it keeps following nonsense-content links endlessly, it's not a human; if it gives up pretty quickly, it might be--this costs/pisses off users but can be used as a test to better train traffic-analysis rules that trigger the other approaches on this list in response to detected likely-crawler traffic). - Model poisoners out of sheer pettiness, if it brings you joy. I also wonder if serving taboo traffic (e.g. legal but beyond-the-pale for most commercial applications porn/erotica) would deter some AI crawlers. There might be front-side content filters that either blacklist or de-prioritize sites whose main content appears (to the crawler) to be at some intersection of inappropriate, prohibited, and not widely-enough related to model output as to be in demand.
- varenc 1y agoAre AI crawlers equipped to get past reCAPTCHA or hCAPTCHA? This seems like exactly the thing these services were meant to stop.
- black_puppydog 1y agoSo the problem is a bunch of AI companies mining our web content for training data without asking and without regard for hosters' effort/bandwidth and the users' service quality. And the proposed remedy is to give them human-labeled data directly in the form of captchas, even more severely degrading the user experience and thus website viability? Color me unconvinced.
- 1oooqooq 1y agoanubis and others allow some user agents to pass without proof of work. bad bots (and user) just use an extension that detect anubis and change the user agent instead. it's well intentioned but just waste electricity from good people in the end. anubis does nothing to impact bad crawlers, well only the laziest ones. but for those generating fake infinite content on the fly is much more efficient.
- pointlessone 1y agoI feel for Codeberg and people who against AI but also I think Anubis can’t die soon enough. It breaks archiving and is very annoying when JS is disabled or when faced with an aggressive ad-blocker. It breaks web in more ways than one.
- TuxPowered 1y agoMaking my web resources IPv6-only has solved the problem for me. I don’t consider this a solution for ever, but for now it’s apparently way too modern or complicated for the A-so-called-I companies.
- speckx 1y agoIn my experience managing a number of IPv6-only sites for clients, they still get crawled and abused, and this goes back years. If anything, it has gotten worse now with all the LLM/AI nonsense.
- anovikov 1y agoHow long it will be until there will be nothing left to scrape and this activity ends? I mean, at some point payback from it will become less than 0 because they will be consuming more AI generated stuff than real (even adjusted for their ability to detect and filter it), and the content itself will be of very little value (because all high value one will be already scraped). Anyone tried to estimate how long that might take?
- nialv7 1y agoIt is just sad we are in a time where measures like Anubis is necessary. The author's efforts are admirable, so I don't mean this personally: but Anubis is a bad product IMHO. It doesn't quite do what it is advertised to do, as evidenced by this post; and it degrades user experience for everybody. And it also stops the website from being indexed by search engines (unless specifically configured otherwise). For example, gitlab.freedesktop.org pages have just disappeared from Google. We need to find a better way.
- ivanstepanovftw 1y agoWhat do these crawlers gather? Just make this data accessible via API calls or direct database download, like Wikipedia did (https://en.wikipedia.org/wiki/Wikipedia:Database_download https://en.wikipedia.org/wiki/Wikipedia:Database_download).
- mmis1000 1y agoThe whole reason of anubis is the bot don't make a damn shit about whether whole data is accessible or not, and even crawl dynamic links in robots.txt in high frequency. Even wikipedia begged for those damn bot about stopping doing this, the data is already accessible in archive here.