4 ms·
Who does Anubis actually stop?
- satvikpendem 2mo agoExactly, it's the same as Cloudflare captchas where only certain blessed devices and browsers can seem to actually pass it. Ironically, adding Anubis accelerates the death of the open web.
- grim_io 2mo agoOpen to whom? There can't be a truly open web if the guys with all the resources in the world have the incentive to absolutely crush you by draining all of your resources. They won't crush you out of malice, but by accident, like an ant.
- nitwit005 2mo agoCloudflare got it's early business from sites being taken down by DDOS attacks. The early users of Anubis were often sites taken down by scrapers. Not a fan of either solution, but the sites simply going offline does not promote an "open web" either.
- mike_hock 2mo agoNo, this is not the same as Cloudflare fascism. You're free to use any JS runtime to solve the challenge.
- pibaker 2mo agoA world without cloudflare does not necessarily mean a world where anyone can access any website. It could very well be a world where anyone can point a DDOS at any website and make it inaccessible to everyone.
- dlcarrier 2mo agoIt's getting to the point that the only way to get past the bot filters is to have bots do everything.
- lrem 2mo agoUh wait, a cookie that can be reused for a time to fetch the rest of the content? That's the exact opposite of what I'd intuitively want for this system...
- recursivecaveat 2mo agoIf you behave yourself and reuse the same cookie/IP instead of trying to hide your identity, other tools can be used to block you if your request volume is crazy. It also means regular people only have to solve once.
- jdlshore 2mo agoHmm, I don’t think I agree. The author is claiming that Anubis is meant to stop individuals using LLMs, but I don’t think that’s its purpose. I believe its purpose is to reduce mass scraping of data that puts excessive load on systems. It does that by increasing the cost of scraping, not by preventing it entirely. Specifically, bad actors were ignoring robots.txt and rotating IPs to make blocking difficult. Anubis serves to make new connections more expensive, so that people will reuse a connection/cookie, and then falls back to normal means of preventing bad actors. (That mass scraping is the result of AI training companies, yes, but it’s the mass scraping that’s the problem, not the LLMs.)
- inigyou 2mo agoIt does it by blocking the scrapers because they are really stupid scrapers. go-away blocks them too, by seeing if they load images, a much quicker test.
- TomatoCo 2mo agoThere's plenty of arguments that mass scrapers have compute to spare but it seems to me that if Anubis makes it 100x more expensive to scrape then, for any given scraping budget, that means you get scraped 100x less. Which is the difference between your server buckling under the load or continuing to serve reliably.
- dullcrisp 2mo agoIf compute isn’t the bottleneck then it wouldn’t make it 100x more expensive to scrape. Of course if it does, then it’ll be effective.
- VladVladikoff 2mo agoIn my experience, not really. Each asset is made by a new proxy, eg some TV somewhere, they send every new request via a new IP address, which is a new machine, and it doesn’t matter how long the response takes, they have already rotated to the next ip immediately and sent another request. Most of these providers have millions of IPs, and it generally doesn’t take millions of requests to scrape a website (unless it’s really big!)
- cjd8 2mo agoFunny, I'm working on a simple tool that pulls the atom feed of the latest patches from lore with Python, and I just slip a User-Agent header into the requests.get() and it works great.
- yellow_lead 2mo ago> The exact adversary Anubis targets defeats it trivially. Wrong, anubis stops mass crawling of web pages, by requiring a proof of work, which makes accessing these websites more expensive. Anubis was not built to stop individual users with llms.
- drum55 2mo agoClaude Opus can make an optimized version of the proof of work solver that’s more than 10000x faster than the javascript one, in about 5 minutes time. Who is this stopping exactly? It’s the wrong tool in the wrong place.
- Macha 2mo agoPractically, the people indiscriminately scraping don't bother to do the work to bypass it or implement the POW test, which results in reduced CPU load for all the properties that were having trouble with scrapers before. Until scrapers start implementing it en masse then, it still serves its purpose.
- greyface- 2mo ago> Until scrapers start implementing it en masse If it becomes widespread (as it has been doing), they will. Anubis' strategy only works while it remains a niche approach only adopted by a small number of sites.
- Macha 2mo agoAnd then people will move to a new solution. I think this is a pragmatist vs idealist debate. The pragmatic answer is that Anubis solves a problem now, and so it will be used until either it doesn't or a better solution presents itself. The idealist approach is that it's obvious that there's ways to get around Anubis, and so some people argue from there that it shouldn't be used. But the only other alternatives being offered are to either eat up the costs (in server resources or engineering time), or to go behind cloudflare with its own tradeoffs.
- novafunc 2mo agoAnubis's primary goal is to prevent web scrapers from DDoSing a website. It's not meant to be an unbeatable challenge or only allow humans like Google's more privacy-invasive captchas. You do the proof of work, you get the content. Not all web scrapers are willing to do the work, which reduces the strain put on web servers. It's by no means a perfect system. It's goals in part prevent it from doing so. It tries to not be too annoying for humans, to not block real users, and not be privacy invasive.
- what 2mo ago>tries not to be annoying The little anime girl is pretty off putting. I bounce when I see it.
- arn3n 2mo agoIt’s meant to be both funny AND highly unprofessional; Anubis makes money of licensing a version of the firewall where you can change the image. It’s a good strategy; personal websites and blogs can display the anime girl without fear, and companies that care about their image end up paying. Win-win.
- ssl-3 2mo agoIt appears in places where that are neither personal websites nor blogs, and that are places where professionals conduct work. For instance: The act of searching the Arch Linux wiki produces a picture of the anime girl. (Should I just not use Arch professionally?)
- jdlshore 2mo agoIf you care that much, you can donate money to Arch to buy a commercial license, or inform your sales rep that you find their conduct unprofessional. Or, yes, you can take the presence of the free version as a sign that the professional service is freeloading, and take your business elsewhere.
- throw0101d 2mo agoI was curious about the name: > Anubis is a Web AI Firewall Utility that weighs the soul of your connection[1] using one or more challenges in order to protect upstream resources from scraper bots. * https://anubis.techaro.lol/docs/ https://anubis.techaro.lol/docs/ > The Weighing of the Heart would take place in Duat (the Underworld), in which the dead were judged by Anubis, using a feather, representing Ma'at, the goddess of truth and justice responsible for maintaining order in the universe. The heart was the seat of the life-spirit (ka). Hearts heavier than the feather of Ma'at were rejected and eaten by Ammit, the Devourer of Souls. * https://en.wikipedia.org/wiki/Weighing_of_souls#Ancient_Egyptian_religion https://en.wikipedia.org/wiki/Weighing_of_souls#Ancient_Egyp...
- aboardRat4 2mo ago>>I was curious about the name: That's knowledge usually learnt in primary school.
- ThrowawayTestr 2mo agoI mean I know about the heart weighing thing but I'm pretty sure I didn't learn it in school.
- theshackleford 2mo ago> That's knowledge usually learnt in primary school. Perhaps where you reside, I’m unsure why you would believe it to be universal.
- j-bos 2mo agoOr scholastic book fairs (showing my age)
- Terr_ 2mo agoGrades 1-5 is a little hyperbolic, IIRC my "World Religions" class was probably middle/junior-high school. In any case, it's not the type of lifetime common-knowledge which warrants your scornful response.
- Rendello 2mo ago
- nitwit005 2mo agoThis mistakes the goal. It's not to block the scrapers, but to discourage excessive (and costly) scraping. The Anubis cookies are bound to particular IP address. The scrapers are often using a large set of IP addresses, so they'll be paying a far higher cost than this suggests.
- ssl-3 2mo agoThere are multiple goals at play. They're easy to find in HN comment sections whenever these topics arise. One goal is to reduce excessive scraping, usually for monetary or performance reasons. This is the goal you mention, and is motivated by a desire keeping the thing working at all. Another goal is to stop bots from ingesting the content, carte blanche. This goal is motivated by a desire to dictate how bots (and by extension, people) may use the information that is otherwise freely-available on the web. These are not the same goals.
- cedws 2mo agoThe cost is nothing. The browser implementation is too slow, mobile devices are too slow, and hash algorithms with hardware acceleration are too fast. You can’t balance these three constraints in a way that only keeps out the bad guys. The hashing is just elaborate obfuscation. Anubis uses SHA256 which isn’t ASIC-resistant, and thanks to Bitcoin you could probably buy one off the shelf.
- eqvinox 2mo agoand yet, it works.
- cedws 2mo agoThe hashing has nothing to do with it. It's just an arbitrary hoop for clients to jump through that filters out the bots that haven't implemented that hoop. So why not just use the client's fingerprint and call it a day? Use their canvas or JA3 fingerprint and I'm willing to bet it would be just as effective.
- patchtopic 2mo agoperhaps the author, instead of the "I'm so clever" theoretical arguments from the client side, actually implemented Anubis on the server side and observe the results. I have set it up on a few sites being relentlessly hammered by clearly idiotic bot traffic, and it drops bot the traffic levels from insane to manageable. I don't even care if the traffic is AI or bots, if the bot has gone to the same effort as the author has the bot may even just access the site in a responsible manner and that's fine.
- PunchyHamster 2mo agoWell, my call on this thing being useless waste of time of everyone involved was correct. Compute wasted on AI tokens alone is probably far more than some cpu for token solving, better invest time into caching it well (or blocking agentic traffic entirely if that's your jam)
- inigyou 2mo agoWe don't know who is doing the really dumb global scraping attack, but that's who it's meant to stop.
- m463 2mo agoAnubis does not stop me from browsing a site (not scraping, browsing with a human at the wheel) On the other hand, Cloudflare and the others stop me dead. "enable javascript and cookies." Ok. "your browser is too old". (I have an old OS with the newest firefox esr that supports it). sigh.
- nonamesleft 2mo agoAnubis also forces cookies and javascript, which deters me from many of the sites that use it that require neither.
- object-a 2mo agoMaybe free market principles apply here: if Anubis fails to reduce scraper/bot load on servers, or blocks too many desired users, then the sites that adopt it would probably scrap it. If they’re keeping it even after all these posts, it must be stopping _some_ sort of undesirable traffic without costing too much desirable traffic.
- seba_dos1 2mo agoThe most interesting property of Anubis is how much content like this that's comically missing the point it inspires.