52 ms·
Why are anime catgirls blocking my access to the Linux kernel?
- jchw 1y ago> This… makes no sense to me. Almost by definition, an AI vendor will have a datacenter full of compute capacity. It feels like this solution has the problem backwards, effectively only limiting access to those without resources or trying to conserve them. A lot of these bots consume a shit load of resources specifically because they don't handle cookies, which causes some software (in my experience, notably phpBB) to consume a lot of resources. (Why phpBB here? Because it always creates a new session when you visit with no cookies. And sessions have to be stored in the database. Surprise!) Forcing the bots to store cookies to be able to reasonably access a service actually fixes this problem altogether. Secondly, Anubis specifically targets bots that try to blend in with human traffic. Bots that don't try to blend in with humans are basically ignored and out-of-scope. Most malicious bots don't want to be targeted, so they want to blend in... so they kind of have to deal with this. If they want to avoid the Anubis challenge, they have to essentially identify themselves. If not, they have to solve it. Finally... If bots really want to durably be able to pass Anubis challenges, they pretty much have no choice but to run the arbitrary code. Anything else would be a pretty straight-forward cat and mouse game. And, that means that being able to accelerate the challenge response is a non-starter: if they really want to pass it, and not appear like a bot, the path of least resistance is to simply run a browser. That's a big hurdle and definitely does increase the complexity of scraping the Internet. It increases more the more sites that use this sort of challenge system. While the scrapers have more resources, tools like Anubis scale the resources required a lot more for scraping operations than it does a specific random visitor. To me, the most important point is that it only fights bot traffic that intentionally tries to blend in. That's why it's OK that the proof-of-work challenge is relatively weak: the point is that it's non-trivial and can't be ignored, not that it's particularly expensive to compute. If bots want to avoid the challenge, they can always identify themselves. Of course, then they can also readily be blocked, which is exactly what they want to avoid. In the long term, I think the success of this class of tools will stem from two things: 1. Anti-botting improvements, particularly in the ability to punish badly behaved bots, and possibly share reputation information across sites. 2. Diversity of implementations. More implementations of this concept will make it harder for bots to just hardcode fastpath challenge response implementations and force them to actually run the code in order to pass the challenge. I haven't kept up with the developments too closely, but as silly as it seems I really do think this is a good idea. Whether it holds up as the metagame evolves is anyone's guess, but there's actually a lot of directions it could be taken to make it more effective without ruining it for everyone.
- o11c 1y ago> A lot of these bots consume a shit load of resources specifically because they don't handle cookies, which causes some software (in my experience, notably phpBB) to consume a lot of resources. (Why phpBB here? Because it always creates a new session when you visit with no cookies. And sessions have to be stored in the database. Surprise!) Forcing the bots to store cookies to be able to reasonably access a service actually fixes this problem altogether. ... has phpbb not heard of the old "only create the session on the second visit, if the cookie was successfully created" trick?
- jchw 1y agophpBB supports browsers that don't support or accept cookies: if you don't have a cookie, the URL for all links and forms will have the session ID in it. Which would be great, but it seems like these bots are not picking those up either for whatever reason.
- MikeDVB 1y agoWe have been seeing our clients' sites being absolutely *hammered* by AI bots trying to blend in. Some of the bots use invalid user agents - they _look_ valid on the surface, but under the slightest scrutiny, it becomes obvious they're not real browsers. Personally I have no issues with AI bots, that properly identify themselves, from scraping content as if the site operator doesn't want it to happen they can easily block the offending bot(s). We built our own proof-of-work challenge that we enable on client sites/accounts as they come under 'attack' and it has been incredible how effective it is. That said I do think it is only a matter of time before the tactics change and these "malicious" AI bots are adapted to look more human / like real browsers. I mean honestly it wouldn't be _that_ hard to enable them to run javascript or to emulate a real/accurate User-Agent. That said they could even run headless versions of the browser engines... It's definitely going to be cat-and-mouse. The most brutal honest truth is that if they throttled themselves as not to totally crash whatever site they're trying to scrape we'd probably have never noticed or gone through the trouble of writing our own proof-of-work challenge. Unfortunately those writing/maintaining these AI bots that hammer sites to death probably either have no concept of the damage it can do or they don't care.
- jimmaswell 1y agoWhat exactly is so bad about AI crawlers compared to Google or Bing? Is there more volume or is it just "I don't like AI"?
- Philpax 1y agoVolume, primarily - the scrapers are running full-tilt, which many dynamic websites aren't designed to handle: https://pod.geraspora.de/posts/17342163 https://pod.geraspora.de/posts/17342163
- immibis 1y agoWhy haven't they been sued and jailed for DDoS, which is a felony?
- ranger_danger 1y agoCriminal convictions in the US require a standard of proof that is "beyond a reasonable doubt" and I suspect cases like this would not pass the required mens rea test, as, in their minds at least (and probably a judge's), there was no ill intent to cause a denial of service... and trying to argue otherwise based on any technical reasoning (e.g. "most servers cannot handle this load and they somehow knew it") is IMO unlikely to sway the court... especially considering web scraping has already been ruled legal, and that a ToS clause against that cannot be legally enforced.
- slowmovintarget 1y agoI thought only capital crimes (murder, for example) held the standard of beyond a reasonable doubt. Lesser crimes require the standard of either a "Preponderance of Evidence" or "Clear and Convincing Evidence" as burden of proof. Still, even by those lesser standards, it's hard to build a case.
- eurleif 1y agoNo, all criminal convictions require proof beyond a reasonable doubt: https://constitution.congress.gov/browse/essay/amdt14-S1-5-5-5/ALDE_00013763/ https://constitution.congress.gov/browse/essay/amdt14-S1-5-5... >Absent a guilty plea, the Due Process Clause requires proof beyond a reasonable doubt before a person may be convicted of a crime.
- _knzv 1y ago[dead]
- PaulHoule 1y ago[flagged]
- dathinab 1y agoyou are overthinking it's a simple as having a nice picture there make this whole thing feel nicer, and give it a bit of personality so you put in some picture/art you like that's it similar any site sing it can change that picture, but there isn't any fundamental problem with the picture, so most can't care to change it
- lxgr 1y ago> This isn’t perfect of course, we can debate the accessibility tradeoffs and weaknesses, but conceptually the idea makes some sense. It was arguably never a great idea to begin with, and stopped making sense entirely with the advent of generative AI.
- oblio 1y agoWhy?
- yuumei 1y ago> The CAPTCHA forces vistors to solve a problem designed to be very difficult for computers but trivial for humans. > Anubis – confusingly – inverts this idea. Not really, AI easily automates traditional captchas now. At least this one does not need extensions to bypass.
- deleted 1y ago[deleted]
- Philpax 1y agoThe argument isn't that it's difficult for them to circumvent - it's not - but that it adds enough friction to force them to rethink how they're scraping at scale and/or self-throttle. I personally don't care about the act of scraping itself, but the volume of scraping traffic has forced administrators' hands here. I suspect we'd be seeing far fewer deployments if the scrapers behaved themselves to begin with.
- davidclark 1y agoThe OP author shows that the cost to scrape an Anubis site is essentially zero since it is a fairly simple PoW algorithm that the scraper can easily solve. It adds basically no compute time or cost for a crawler run out of a data center. How does that force rethinking?
- hooverd 1y agoThe problem with crawlers if that they're functionally indistinguishable from your average malware botnet in behavior. If you saw a bunch of traffic from residential IPs using the same token that's a big tell.
- Philpax 1y agoThe cookie will be invalidated if shared between IPs, and it's my understanding that most Anubis deployments are paired with per-IP rate limits, which should reduce the amount of overall volume by limiting how many independent requests can be made at any given time. That being said, I agree with you that there are ways around this for a dedicated adversary, and that it's unlikely to be a long-term solution as-is. My hope is that the act of having to circumvent Anubis at scale will prompt some introspection (do you really need to be rescraping every website constantly?), but that's hopeful thinking.
- yborg 1y ago>do you really need to be rescraping every website constantly Yes, because if you believe you out-resource your competition, by doing this you deny them training material.
- anotherhue 1y agoSurely the difficulty factor scales with the system load?
- lousken 1y agoaren't you happy? at least you see catgirl
- jayrwren 1y agoliterally the top link when I search for his exact text "why are anime catgirls blocking my access to the Linux kernel?" https://lock.cmpxchg8b.com/anubis.html https://lock.cmpxchg8b.com/anubis.html Maybe travis needs more google-fu. maybe that includes using duckduckgo?
- Macha 1y agoThe top link when you search the title of the article is the article itself? I am shocked, shocked I say.
- ksymph 1y agoThis is neither here nor there but the character isn't a cat. It's in the name, Anubis, who is an Egyptian deity typically depicted as a jackal or generic canine, and the gatekeeper of the afterlife who weighs the souls of the dead (hence the tagline). So more of a dog-girl, or jackal-girl if you want to be technical.
- pak9rabid 1y agoWell, thank you for that. That's a great weight off me mind.
- JdeBP 1y ago... but entirely lacking the primary visual feature that Anubis had.
- esperent 1y agoEvery representation I've ever seen of Anubis - including remarkably well preserved statues from antiquity - are either a male human body with a canine head, or fully canine. This anime girl is not Anubis. It's a modern cartoon characters that simply borrows the name because it sounds cool, without caring anything about the history or meaning behind it. Anime culture does this all the time, drawing on inspiration from all cultures but nearly always only paying the barest lip service to the original meaning. I don't have an issue with that, personally. All cultures and religions should be fair game as inspiration for any kind of art. But I do have an issue with claiming that the newly inspired creation is equivalent in any way to the original source just because they share a name and some other very superficial characteristics.
- qwery 1y agoI think you're taking it a bit too seriously. In turn, I am, of course, also taking it too seriously. > I do have an issue with claiming that the newly inspired creation is equivalent in any way to the original source Nobody is claiming that the drawing is Anubis or even a depiction of Anubis, like the statues etc. you are interested in. It's a mascot. "Mascot design by CELPHASE" -- it says, in the screenshot. Generally speaking -- I can't say that this is what happened with this project -- you would commission someone to draw or otherwise create a mascot character for something after the primary ideation phase of the something. This Anubis-inspired mascot is, presumably, Anubis-inspired because the project is called Anubis, which is a name with fairly obvious connections to and an understanding of "the original source". > Anime culture does this all the time, ... I don't know what bone you're picking here. This seems like a weird thing to say. I mean, what anime culture? It's a drawing on a website. Yes, I can see the manga/anime influence -- it's a very popular, mainstream artform around the world.
- rnhmjoj 1y agoI don't understand, why do people resort to this tool instead of simply blocking by UA string or IP address. Are there so many people running these AI crawlers? I blackholed some IP blocks of OpenAI, Mistral and another handful of companies and 100% of this crap traffic to my webserver disappeared.
- hooverd 1y agoless savory crawlers use residential proxies and are indistinguishable from malware traffic
- WesolyKubeczek 1y agoYou should read more. AI companies use residential proxies and mask their user agents with legitimate browser ones, so good luck blocking that.
- rnhmjoj 1y agoWhich companies are we talking about here? In my case the traffic was similar to what was reported here[1]: these are crawlers from Google, OpenAI, Amazon, etc. they are really idiotic in behaviour, but at least report themselves correctly. [1]: https://pod.geraspora.de/posts/17342163 https://pod.geraspora.de/posts/17342163
- nemothekid 1y agoOpenAI/Anthropic/Perplexity aren't the bad actors here. If they are, they are relatively simply to block - why would you implement an Anubis PoW MITM Proxy, when you could just simply block on UA? I get the sense many of the bad actors are simply poor copycats that are poorly building LLMs and are scraping the entire web without a care in the world
- rnhmjoj 1y ago> why would you implement an Anubis PoW MITM Proxy, when you could just simply block on UA? That's in fact what I was asking: I've only seen traffic from these kind of companies and I've easily blocked them without an annoying PoW scheme. I have yet to see any of these bad actors and I'm interested in knowing who they actually are.
- WesolyKubeczek 1y agoI disagree with the post author in their premise that things like Anubis are easy to bypass if you craft your bot well enough and throw the compute at it. Thing is, the actual lived experience of webmasters tells that the bots that scrape the internets for LLMs are nothing like crafted software. They are more like your neighborhood shit-for-brain meth junkies competing with one another who makes more robberies in a day, no matter the profit. Those bots are extremely stupid. They are worse than script kiddies’ exploit searching software. They keep banging the pages without regard to how often, if ever, they change. If they were 1/10th like many scraping companies’ software, they wouldn’t be a problem in the first place. Since these bots are so dumb, anything that is going to slow them down or stop them in their tracks is a good thing. Short of drone strikes on data centers or accidents involving owners of those companies that provide networks of botware and residential proxies for LLM companies, it seems fairly effective, doesn’t it?
- fluoridation 1y agoHmm... What if instead of using plain SHA-256 it was a dynamically tweaked hash function that forced the client to run it in JS?
- VMG 1y agocrawlers can run JS, and also invest into running the Proof-Of-JS better than you can
- fluoridation 1y agoIf we're presupposing an adversary with infinite money then there's no solution. One may as well just take the site offline. The point is to spend effort in such a way that the adversary has to spend much more effort, hopefully so much it's impractical.
- tjhorner 1y agoAnubis doesn't target crawlers which run JS (or those which use a headless browser, etc.) It's meant to block the low-effort crawlers that tend to make up large swaths of spam traffic. One can argue about the efficacy of this approach, but those higher-effort crawlers are out of scope for the project.
- Imustaskforhelp 1y agoreminds of how wikipedia literally has all the data available even in a nice format just for scrapers (I think) and even THEN, there are some scrapers which still scraped wikipedia and actually made wikipedia lose some money so much that I am pretty sure that some official statement had to be made or they disclosed about it without official statement. Even then, man I feel like you yourself can save on so many resources (both yours) and (wikipedia) if scrapers had the sense to not scrape wikipedia and instead follow wikipedia's rules
- scratchyone 1y agowait but then why bother with this PoW system at all? if they're just trying to block anyone without JS that's way easier and doesn't require slowing things down for end users on old devices.
- ksymph 1y agoReading the original release post for Anubis [0], it seems like it operates mainly on the assumption that AI scrapers have limited support for JS, particularly modern features. At its core it's security through obscurity; I suspect that as usage of Anubis grows, more scrapers will deliberately implement the features needed to bypass it. That doesn't necessarily mean it's useless, but it also isn't really meant to block scrapers in the way TFA expects it to. [0] https://xeiaso.net/blog/2025/anubis/ https://xeiaso.net/blog/2025/anubis/
- jhanschoo 1y agoYour link explicitly says: > It's a reverse proxy that requires browsers and bots to solve a proof-of-work challenge before they can access your site, just like Hashcash. It's meant to rate-limit accesses by requiring client-side compute light enough for legitimate human users and responsible crawlers in order to access but taxing enough to cost indiscriminate crawlers that request host resources excessively. It indeed mentions that lighter crawlers do not implement the right functionality in order to execute the JS, but that's not the main reason why it is thought to be sensible. It's a challenge saying that you need to want the content bad enough to spend the amount of compute an individual typically has on hand in order to get me to do the work to serve you.
- ksymph 1y agoHere's a more relevant quote from the link: > Anubis is a man-in-the-middle HTTP proxy that requires clients to either solve or have solved a proof-of-work challenge before they can access the site. This is a very simple way to block the most common AI scrapers because they are not able to execute JavaScript to solve the challenge. The scrapers that can execute JavaScript usually don't support the modern JavaScript features that Anubis requires. In case a scraper is dedicated enough to solve the challenge, Anubis lets them through because at that point they are functionally a browser. As the article notes, the work required is negligible, and as the linked post notes, that's by design. Wasting scraper compute is part of the picture to be sure, but not really its primary utility.
- 1y ago
- iefbr14 1y agoI wouldn't be surprised if just delaying the server response by some 3 seconds will have the same effect on those scrapers as Anubis claims.
- ranger_danger 1y agoYea I'm not convinced unless somehow the vast majority of scrapers aren't already using headless browsers (which I assume they are). I feel like all this does is warm the planet.
- kingstnap 1y agoThere is literally no point wasting 3 seconds of a computer's time and it's expensive wasting 3 seconds of a person's time. That is literally an anti-human filter.
- Imustaskforhelp 1y agoFrom tjhorner on this same thread "Anubis doesn't target crawlers which run JS (or those which use a headless browser, etc.) It's meant to block the low-effort crawlers that tend to make up large swaths of spam traffic. One can argue about the efficacy of this approach, but those higher-effort crawlers are out of scope for the project." So its meant/preferred to block low effort crawlers which can still cause damage if you don't deal with them. a 3 second deterrent seems good in that regard. Maybe the 3 second deterrent can come as in rate limiting an ip? but they might use swath's of ip :/
- OkayPhysicist 1y agoAnubis exists specifically to handle the problem of bots dodging IP rate limiting. The challenge is tied to your IP, so if you're cycling IPs with every request, you pay dramatically more PoW than someone using a single IP. It's intended to be used in depth with IP rate limiting.
- loeg 1y agoAnubis easily wastes 3 seconds of a human's time already.
- xena 1y ago[dead]
- withinrafael 1y agoThe security policy that didn't exist until a few hours ago?
- david_allison 1y agoAdded on March 18: https://github.com/TecharoHQ/.github/commits/main/SECURITY.md https://github.com/TecharoHQ/.github/commits/main/SECURITY.m... Copied to the root of the repo after the disclosure ref: https://github.com/TecharoHQ/anubis/issues/1002#issuecomment-3207867466 https://github.com/TecharoHQ/anubis/issues/1002#issuecomment...
- Borgz 1y agoIn a different repository, though. I think it's understandable that someone would miss it.
- withinrafael 1y agoAdding a security policy to an unrelated repository is easily missed and questionably applicable.
- tptacek 1y agoYou needed to have a security contact on your website, or at least in the repo. You did not. You assumed security researchers would instead back out to your Github account's repository list, find the .github repository, and look for a security policy there. That's not a thing! I'm really surprised you wrote this.
- qualeed 1y ago>I'm really surprised you wrote this. I agree with the rest of your comment, but this seems like a weird little jab to add on for no particular reason. Am I misinterpreting?
- immibis 1y agoThe actual answer to how this blocks AI crawlers is that they just don't bother to solve the challenge. Once they do bother solving the challenge, the challenge will presumably be changed to a different one.
- Arnavion 1y ago>This dance to get access is just a minor annoyance for me, but I question how it proves I’m not a bot. These steps can be trivially and cheaply automated. >I think the end result is just an internet resource I need is a little harder to access, and we have to waste a small amount of energy. No need to mimic the actual challenge process. Just change your user agent to not have "Mozilla" in it; Anubis only serves you the challenge if it has that. For myself I just made a sideloaded browser extension to override the UA header for the handful of websites I visit that use Anubis, including those two kernel.org domains. (Why do I do it? For most of them I don't enable JS or cookies for so the challenge wouldn't pass anyway. For the ones that I do enable JS or cookies for, various self-hosted gitlab instances, I don't consent to my electricity being used for this any more than if it was mining Monero or something.)
- throw84a747b4 1y ago[flagged]
- gruez 1y ago>Not only is Anubis a poorly thought out solution from an AI sympathizer [...] But the project description describes it as a project to stop AI crawlers? > Weighs the soul of incoming HTTP requests to stop AI crawlers
- throw84a747b4 1y agoWhy would a company that wants to stop AI crawlers give talks on LLMs and diffusion models at AI conferences? Why would they use AI art for the first Anubis mascot until GitHub users called out the hypocrisy on the issue tracker? Why would they use Stable Diffusion art in their blogposts until Mastodon and Bluesky users called them out on it?
- Imustaskforhelp 1y agoI am not again AI art completely since I think of it as an editing instead of art itself. My thoughts on AI art are nuanced and worth discussing some other day, lets talk about the author of anubis/story of anubis So, I hope you know the entire story behind Anubis, firstly they were hosting their own git server (I think?) and amazon's ai related department was basically ddosing their server in some sense by trying to scrape it and they created anubis in a way to prevent that. The idea isn't that new, it is just proof of work and they created it firstly for their own use and I think that they are An AI researcher/ related to AI, so for them using AI pics wasn't that big of a deal and pretty sure that they had some reason behind it and even that has been changed. Stop whining about free projects/labour man. The same people comment oh well these AI scrapers are scraping so many websites and taking livelihood of website makers and now you have someone who just gave it to ya for free and you are nitpicking the wrong things. You can just fork it without the anime images or without the AI thing if you don't align with them and their philosophy. Man now I feel the mandela effect as I read it somewhere on their blog or any thing that they themselves feel the hypocrisy or something along that (pardon me if I am wrong, I usually am) But they themselves (I think?) would like to get rid of working in the AI industry while making anti AI scraper but they might need more donations iirc and they themselves know the hypocrisy.
- valiant55 1y agoI really don't understand the hostility towards the mascot. I can't think of a bigger red flag.
- Borgz 1y agoFunny to say this when the article literally says "nothing wrong with mascots!" Out of curiosity, what did you read as hostility?
- valiant55 1y agoOh I totally reacted to the title. The last few times Anubis has been the topic there's always comments about "cringy" mascot and putting that front and center in the title just made me believe that anime catgirls was meant as an insult.
- Imustaskforhelp 1y agoHonestly I am okay with anime catgirls since I just find it funny but still it would be cool to see linux related stuff. Imagine mr tux penguin gif of him racing in like supertuxcart for the linux website. sourcehut also uses anubis but they have removed the anime catgirl thing with their own logo, I think disroot also does that I am not sure though
- Arnavion 1y agoSourcehut uses go-away, not Anubis.
- Imustaskforhelp 1y agohttps://sourcehut.org/blog/2025-04-15-you-cannot-have-our-users-data/ https://sourcehut.org/blog/2025-04-15-you-cannot-have-our-us... > As you may have noticed, SourceHut has deployed Anubis to parts of our services to protect ourselves from aggressive LLM crawlers. Its nice that sourcehut themselves have talked about it on their own blog but I had discovered this through the anubis website themselves showcases or soemthing like that iirc.
- jmclnx 1y ago>The CAPTCHA forces vistors to solve a problem designed to be very difficult for computers but trivial for humans Not for me, I have nothing but a hard time solving CAPTCHAs, ahout 50% of the time I give up after 2 tries.
- serf 1y agoit's still certainly trivial for you compared to mentally computing a SHA256 op.
- johnea 1y agoMy biggest bitch is that it requires JS and cookies... Although the long term problem is the business model of servers paying for all network bandwidth. Actual human users have consumed a minority of total net bandwidth for decades: https://www.atom.com/blog/internet-statistics/ https://www.atom.com/blog/internet-statistics/ Part 4 shows bots out using humans in 1996 8-/ What are "bots"? This needs to include goggleadservices, PIA sharing for profit, real-time ad auctions, and other "non-user" traffic. The difference between that and the LLM training data scraping, is that the previous non-human traffic was assumed, by site servers, to increase their human traffic, through search engine ranking, and thus their revenue. However the current training data scraping is likely to have the opposite effect: capturing traffic with LLM summaries, instead of redirecting it to original source sites. This is the first major disruption to the internet's model of finance since ad revenue look over after the dot bomb. So far, it's in the same category as the environmental disaster in progress, ownership is refusing to acknowledge the problem, and insisting on business as usual. Rational predictions are that it's not going to end well...
- jerf 1y ago"Although the long term problem is the business model of servers paying for all network bandwidth." Servers do not "pay for all the network bandwidth" as if they are somehow being targeted for fees and carrying water for the clients that are somehow getting it for "free". Everyone pays for the bandwidth they use, clients, servers, and all the networks in between, one way or another. Nobody out there gets free bandwidth at scale. The AI scrapers are paying lots of money to scrape the internet at the scales they do.
- Imustaskforhelp 1y agoThe Ai scrapers are most likely vc funded and all they care about is getting as much data as possible and not worry about the costs. They are hiring machines at scale too so definitely bandwidth etc. are cheaper for them too. Maybe use a provider that doesn't have too much bandwidth issues (hetzner?) But still, the point being that you might be hosting website on your small server and that scraper with its machines beast can come and effectively ddos your server looking for data to scrape. Deterring them is what matters so that the economical scale finally slide back to our favours again.
- superkuh 1y agoKernel.org* just has to actually configure Anubis rather than deploying the default broken config. Enable the meta-refresh proof of work rather than relying on the corporate browsers only bleeding edge javascript application proof of work. * or whatever site the author is talking about, his site is currently inaccessible due to the amount of people trying to load it.
- bogwog 1y agoI wonder if the best solution is still just to create link mazes with garbage text like this: https://blog.cloudflare.com/ai-labyrinth/ https://blog.cloudflare.com/ai-labyrinth/ It won't stop the crawlers immediately, but it might lead to an overhyped and underwhelming LLM release from a big name company, and force them to reassess their crawling strategy going forward?
- ronsor 1y agoThat won't work, because garbage data is filtered after the full dataset is collected anyway. Every LLM trainer these days knows that curation is key.
- bogwog 1y agoIf the "garbage data" is AI generated, it'll be hard or impossible to filter.
- creatonez 1y agoCrawlers already know how to stop crawling recursive or otherwise excessive/suspicious content. They've dealt with this problem long before LLM-related crawling.
- rootsudo 1y agoWhen I instantly read it, I knew it was anubis. I hope the anime catgirls never disapear from that project :)
- bakugo 1y agoIt's more likely that the project itself will disappear into irrelevance as soon as AI scrapers bother implementing the PoW (which is trivial for them, as the post explains) or figure out that they can simply remove "Mozilla" from their user-agent to bypass it entirely.
- skydhash 1y agoIt's more about the (intentional?) DDoS from AI scrappers, than preventing them from accessing the content. Bandwidth is not cheap.
- dingnuts 1y ago[flagged]
- shkkmo 1y ago> PoW increases the cost for the bots which is great. But not by any meaningful amount as explained in the article. All it actually does is rely on it's obscurity while interfering with legitimate use.
- nialv7 1y ago> Fuck AI scrapers, and fuck all this copyright infringement at scale. Yes, fuck them. Problem is Anubis here is not doing the job. As the article already explains, currently Anubis is not adding a single cent to the AI scrappers' costs. For Anubis to become effective against scrappers, it will necessarily have to become quite annoying for legitimate users.
- Gibbon1 1y agoBest response to AI scrapers is to poison their models.
- leumon 1y agoSeems like ai bots are indeed bypassing the challenge by computing it: https://social.anoxinon.de/@Codeberg/115033790447125787 https://social.anoxinon.de/@Codeberg/115033790447125787
- deleted 1y ago[deleted]
- debugnik 1y agoThat's not bypassing it, that's them finally engaging with the PoW challenge as intended, making crawling slower and more expensive, instead failing to crawl at all, which is more of a plus. This however forces servers to increase the challenge difficulty, which increases the waiting time for the first-time access.
- nialv7 1y agoObviously the developer of Anubis thinks it is bypassing: https://github.com/TecharoHQ/anubis/issues/978 https://github.com/TecharoHQ/anubis/issues/978
- debugnik 1y agoFair, then I obviously think Xe may have a kinda misguided understanding of their own product. I still stand by the concept I stated above.
- rhaps0dy 1y agolatest update from Xe: > After further investigation and communication. This is not a bug. The threat actor group in question installed headless chrome and simply computed the proof of work. I'm just going to submit a default rule that blocks huawei.
- scratchyone 1y agothis kinda proves the entire project doesn't work if they have to resort to manual IP blocking lol
- easterncalculus 1y ago[flagged]
- listic 1y agoKind of. The author is asking nicely not to remove the picture of Anubis. The tool is open source.
- hansjorg 1y agoIf you want a tip my friend, just block all of Huawei Cloud by ASN.
- wging 1y ago... looks like they did: https://github.com/TecharoHQ/anubis/pull/1004 https://github.com/TecharoHQ/anubis/pull/1004, timestamped a few hours after your comment.
- scratchyone 1y agolmfao so that kinda defeats the entire point of this project if they have to resort to a manual IP blocklist anyways
- BLKNSLVR 1y agoI would actually say that it's been successful in determining at least one, so far, large scale abuser, which can the be blocked via more traditional methods. I have my own project that finds malicious traffic IP addresses, and through searching through the results, it's allowed me to identify IP address ranges to be blocked completely. Yielding useful information may not have been what it was designed to do, but it's still a useful outcome. Funny thing about Anubis' viral popularity is that it was designed to just protect the author's personal site from a vast army of resource-sucking marauders, and grew because it was open sourced and a LOT of other people found it useful and effective.
- sandywaffles 1y agoI think that was already common knowledge as hansjorg above suggests
- sugarpimpdorsey 1y agoEvery time I see one of these I think it's a malicious redirect to some pervert-dwelling imageboard. On that note, is kernel.org really using this for free and not the paid version without the anime? Linux Foundation really that desperate for cash after they gas up all the BMW's?
- qualeed 1y agoIt's crazy (especially considering anime is more popular now than ever; netflix alone is making billions a year on anime) that people see a completely innocent little anime picture and immediately think "pervent-dwelling imageboard".
- Seattle3503 1y agoTo be fair, that's the sort of place where I spend most of my free time.
- gruez 1y ago"Anime pfp" stereotype is alive and well.
- turtletontine 1y agoEven if the images aren’t the kind of sexualized (or downright pornographic) content this implies… having cutesy anime girls pop up when a user loads your site is, at best, wildly unprofessional. (Dare I say “cringe”?) For something as serious and legit as kernel.org to have this, I do think it’s frankly shocking and unacceptable.
- deleted 1y ago[deleted]
- antiloper 1y agoIf anime girls prevent LLM scraper sympathizers from interacting with the kernel, that's a good thing and should be encouraged more!
- listic 1y agoSo... Is Anubis actually blocking bots because they didn't bother to circumvent it?
- loloquwowndueo 1y agoBasically. Anubis is meant to block mindless, careless, rude bots with seemingly no technically proficient human behind the process; these bots tend to be very aggressive and make tons of requests bringing sites down. The assumption is that if you’re the operator of these bots and care enough to implement the proof of work challenge for Anubis you could also realize your bot is dumb and make it more polite and considerate. Of course nothing precludes someone implementing the proof of work on the bot but otherwise leaving it the same (rude and abusive). In this case Anubis still works as a somewhat fancy rate limiter which is still good.
- elcritch 1y agoEssentially the Pow aspect is pointless then? They could require almost any arbitrary thing.
- loloquwowndueo 1y agoWhat else do you envision being used instead of proof of work?
- semiquaver 1y agoRot13 a challenge string. It could be any arbitrary function.
- loloquwowndueo 1y agoThat wouldn’t have the fallback rate-limiting functionality. It’s too cheap.
- 1y ago
- zb3 1y agoAnubis doesn't use enough resources to deter AI bots. If you really want to go this way, use React, preferably with more than one UI framework.
- Borg3 1y agoOh, its time to bring Internet back to humans. Maybe its time to treat first layer of Internet just as transport. Then, layer large VPN networks and put services there. People will just VPN to vISP to reach content. Different networks, different interests :) But this time dont fuck up abuse handling. Someone is doing something fishy? Depeer him from network (or his un-cooperating upstream!).
- naikrovek 1y ago[flagged]
- s1mplicissimus 1y agoIsn't using an anime catgirl avatar the exact opposite of "look at meee"?
- naikrovek 1y agono. it's someone wanting attention and feeling ok creating an interstitial page to capture your attention which does not prove you're a human while saying that the page proves you're human. the entire thing is ridiculous. and only those who see no problem shoving anime catgirls into the face of others will deploy it. maybe that's a lot of people; maybe only I object to this. The reality is that there's no technical reason to deploy it, as called out in the linked blog article, so the only reason to do this is a "look at meee" reason, or to announce that one is a fan of this kind of thing, which is another "look at meee"-style reason. Why do I object to things like this? Because once you start doing things like this, doing things for attention, you must continually escalate in order to keep capturing that attention. Ad companies do this, and they don't see a problem in escalation, and they know they have to do it. People quickly learn to ignore ads, so in order to make your page loads count as an advertiser, you must employ means which draw attention to your ads. It's the same with people who no longer draw attention to themselves because they like anime catgirls. Now they must put an interstitial page up to force you to see that they like anime catgirls. We've already established that the interstitial page accomplishes nothing other than showing you the image, so showing the image must be the intent. That is what I object to.
- AuthAuth 1y agoIts just a mascot you are projecting way to much.
- serf 1y agoI don't care that they use anime catgirls. What I do care about is being met with something cutesy in the face of a technical failure anywhere on the net. I hate Amazon's failure pets, I hate google's failure mini-games -- it strikes me as an organizational effort to get really good at failing rather than spending that same effort to avoid failures all together. It's like everyone collectively thought the standard old Apache 404 not found page was too feature-rich and that customers couldn't handle a 3 digit error, so instead we now get a "Whoops! There appears to be an error! :) :eggplant: :heart: :heart: <pet image.png>" and no one knows what the hell is going on even though the user just misplaced a number in the URL.
- JdeBP 1y agoGuru Meditations and Sad Macs are not your thing?
- Hizonner 1y agoThat also got old when you got it again and again while you were trying to actually do something. But there wasn't the space to fit quite as much twee on the screen...
- krige 1y agoFWIW second and third iteration of AmigaOS didn't have "Guru Meditation"; instead it bluntly labeled the numbers as error and task.
- pak9rabid 1y agoI hear this
- xandrius 1y agoThe original versions were a way to make fun even a boring event such as a 404. If the page stops conveying the type of error to the user then it's just bad UX but also vomiting all the internal jargon to a non-tech user is bad UX. So, I don't see an error code + something fun to be that bad. People love dreaming of the 90s wild web and hate the clean cut soulless corp web of today, so I don't see how having fun error pages to be such an issue?
- efilife 1y agoThis cartoon mascot has absolutely nothing to do with anime If you disagree, please say why
- ge96 1y agoOh I saw this recently on ffmpeg's site, pretty fun
- raffraffraff 1y agoHN hug of death
- mr_toad 1y agoI’m getting a black page. Not sure if it’s an ironic meta commentary, or just my ad blocker.
- johnisgood 1y agoI like hashcash. https://github.com/factor/factor/blob/master/extra/hashcash/hashcash.factor https://github.com/factor/factor/blob/master/extra/hashcash/... https://bitcoinwiki.org/wiki/hashcash https://bitcoinwiki.org/wiki/hashcash
- loloquwowndueo 1y agoAnubis is based on hashcash concepts - just adapted to a web request flow. Basically the same thing - moderately expensive for the sender/requester to compute, insanely cheap for the server/recipient to verify.
- littlecranky67 1y agoWe need bitcoin-based lightning nano-payments for such things. Like visiting the website will cost $0.0001 cent, the lightning invoice is embedded in the header and paid for after single-click confirmation or if threshold is under a pre-configured value. Only way to deal with AI crawlers and future AI scams. With the current approach we just waste the energy, if you use bitcoin already mined (=energy previously wasted) it becomes sustainable.
- tonymet 1y agoSo it's a paywall with -- good intentions -- and even more accessibility concerns. Thus accelerating enshittification. Who's managing the network effects? How do site owners control false positives? Do they have support teams granting access? How do we know this is doing any good? It's convoluted security theater mucking up an already bloated , flimsy and sluggish internet. It's frustrating enough to guess schoolbuses every time I want to get work done, now I have to see porfnified kitty waifus (openwrt is another community plagued with this crap)
- tonymet 1y agohere is the community post with Anubis pro / con experiences https://forum.openwrt.org/t/trying-out-anubis-on-the-wiki/232675/44 https://forum.openwrt.org/t/trying-out-anubis-on-the-wiki/23...
- andromaton 1y agoHug of death https://archive.ph/BSh1l https://archive.ph/BSh1l
- xphos 1y agoYeah the PoW is minor for botters but annoying people. I think the only positive is if enough people see anime girls on there screens there might actually be political pressure to make laws against rampent bot crawling
- Havoc 1y ago> PoW is minor for botters But still enough to prevent a billion request DDoS These sites have been search engine scrapped forever. It’s not about blocking bots entirely just about this new wave of fuck you I don’t care if your host goes down quasi malicious scrappers
- st3fan 1y ago"But still enough to prevent a billion request DDoS" - don't you just do the PoW once to get a cookie and then you can browse freely?
- seba_dos1 1y agoYes, but a single bot is not a concern. It's the first "D" in DDoS that makes it hard to handle (and these bots tend to be very, very dumb - which often happens to make them more effective at DDoSing the server, as they're taking the worst and the most expensive ways to scrape content that's openly available more efficiently elsewhere)
- elcritch 1y agoReading TFA, those billions requests would cost web crawlers what about $100 in compute?
- heap_perms 1y ago> I host this blog on a single core 128MB VPS No wonder the site is being hugged to death. 128MB is not a lot. Maybe it's worth to upgrade if you post to hacker news. Just a thought.
- bawolff 1y agoIt doesnt take much to host a static website. Its all the dynamic stuff/frameworks/db/etc that bogs everything down.
- tambourine_man 1y agoStill, 128MB is not enough to even run Debian let alone Apache/NGINX. I’m on my phone, but it doesn’t seem like the author is using Cloudflare or another CDN. I’d like to know what they are doing.
- ronsor 1y ago128MB is more than enough to run Debian and serve a static site. I had no issue with doing it a decade ago and it still works fine. How much memory do you think it actually takes to accept a TLS connection and copy files from disk to a socket?
- tambourine_man 1y agoModern Linux is much less frugal these days: https://wiki.debian.org/DebianEdu/Documentation/Bullseye/Requirements https://wiki.debian.org/DebianEdu/Documentation/Bullseye/Req... * Thin clients with only 256 MiB RAM and 400 MHz are possible, though more RAM and faster processors are recommended. * For workstations, diskless workstations and standalone systems, 1500 MHz and 1024 MiB RAM are the absolute minimum requirements. For running modern webbrowsers and LibreOffice at least 2048 MiB RAM is recommended.
- bawolff 1y agoThat's for some educational distro, which presumably is running some fancy desktop environment with fancy GUI programs. I don't think that is reflective of what a web server needs.
- johnklos 1y agoThis is a usually technical crowd, so I can't help but wonder if many people genuinely don't get it, or if they are just feigning a lack of understanding to be dismissive of Anubis. Sure, the people who make the AI scraper bots are going to figure out how to actually do the work. The point is that they hadn't, and this worked for quite a while. As the botmakers circumvent, new methods of proof-of-notbot will be made available. It's really as simple as that. If a new method comes out and your site is safe for a month or two, great! That's better than dealing with fifty requests a second, wondering if you can block whole netblocks, and if so, which. This is like those simple things on submission forms that ask you what 7 + 2 is. Of course everyone knows that a crawler can calculate that! But it takes a human some time and work to tell the crawler HOW.
- odo1242 1y agoAlso, it forces the crawler to gain code execution capabilities, which for many companies will just make them give up and scrape someone else.
- wredcoll 1y agoI don't know if you've noticed, but there's a few websites these days that use javascript as part of their display logic.
- odo1242 1y agoYes, and those sites take way more effort to crawl than other sites. They may still get crawled, but likely less often than the ones that don't use JavaScript for rendering (which is the main purpose of Anubis - saving bandwidth from crawlers who crawl sites way too often). (Also, note the difference between using JavaScript for display logic and requiring JavaScript to load any content at all. Most websites do the first, the second isn't quite as common.)
- cakealert 1y agoThis arms race will have a terminus. The bots will eventually be indistinguishable from humans. Some already are.
- sidewndr46 1y ago> The CAPTCHA forces vistors to solve a problem designed to be very difficult for computers but trivial for humans I'm an unsure if this deadpan humor or if the author has never tried to solve a CAPTCHA that is something like "select the squares with an orthodox rabbi present"
- bawolff 1y agoWell the problem is that computers got good at basically everything. Early 2000s captchas really were like that.
- bawolff 1y ago> This… makes no sense to me. Almost by definition, an AI vendor will have a datacenter full of compute capacity. It feels like this solution has the problem backwards, effectively only limiting access to those without resources or trying to conserve them. Counterpoint - it seems to work. People use anubis because its the best of bad options. If theory and reality disagree, it means either you are missing something or your theory is wrong.
- semiquaver 1y agoCounter-counter point: it only stopped them for a few weeks and now it doesn’t work: https://news.ycombinator.com/item?id=44914773 https://news.ycombinator.com/item?id=44914773
- Aachen 1y agoOnly Huawei so far, no? That could be easy to block on a network level for the time being Of course we knew from the beginning that this first stage of "bots don't even try to solve it, no matter the difficulty" isn't a forever solution
- jeroenhd 1y agoAliCloud also seems to send a more capable scraper army, but so far they're not using botnets ("residential proxies") to hide their bad practices.
- jeroenhd 1y agoGeoblocking China and Singapore solves that problem, it seems, at least the non-residential IPs (though I also see a lot of aggressive bots coming from residential IP space from China). I wish the old trick of sending CCP-unfriendly content to get the great firewall to kill the connection for you still worked, but in the days of TLS everywhere that doesn't seem to work anymore.
- throwaway984393 1y ago[dead]
- extraduder_ire 1y agoWith the asymmetry of doing the PoW in javascript versus compiled c code, I wonder if this type of rate limiting is ever going to be directly implemented into regular web browsers. (I assume there's already plugins for curl/wget) Other than Safari, mainstream browsers seem to have given up on considering browsing without javascript enabled a valid usecase. So it would purely be a performance improvement thing.
- Aachen 1y agoApple supports people that want to not use their software as the gods at Apple intended it? What parallel universe Version of Apple is this! Seriously though, does anything of Apple's work without JS, like Icloud or Find my phone? Or does Safari somehow support it in a way that other browsers don't?
- extraduder_ire 1y agoLast I checked, safari still had a toggle to disable javascript long after both chrome and firefox removed theirs. That's what I was referring to.
- jonathanyc 1y ago> The idea of “weighing souls” reminded me of another anti-spam solution from the 90s… believe it or not, there was once a company that used poetry to block spam! > Habeas would license short haikus to companies to embed in email headers. They would then aggressively sue anyone who reproduced their poetry without a license. The idea was you can safely deliver any email with their header, because it was too legally risky to use it in spam. Kind of a tangent but learning about this was so fun. I guess it's ultimately a hack for there not being another legally enforceable way to punish people for claiming "this email is not spam"? IANAL so what I'm saying is almost certainly nonsense. But it seems weird that the MIT license has to explicitly say that the licensed software comes with no warranty that it works, but that emails don't have to come with a warranty that they are not spam! Maybe it's hard to define what makes an email spam, but surely it is also hard to define what it means for software to work. Although I suppose spam never e.g. breaks your centrifuge.
- senectus1 1y agothe action is great, anubis is a very clever idea i love it. I'm not a huge fan of the anime thing, but i can live with it.
- qwertytyyuu 1y agoIsn’t animus a dog? So it should be anime dog/wolf girl rather than cat girl?
- Twisol 1y agoYes, Anubis is a dog-headed or jackal-headed god. I actually can't find anywhere on the Anubis website where they talk about their mascot; they just refer to her neutrally as the "default branding". Since dog girls and cat girls in anime can look rather similar (both being mostly human + ears/tail), and the project doesn't address the point outright, we can probably forgive Tavis for assuming catgirl.
- ok123456 1y agoWhy is kernel.org doing this for essentially static content? Cache control headers and ETAGS should solve this. Also, the Linux kernel has solved the C10K problem.
- mixologic 1y agoBecause its static content that is almost never cached because its infrequently accessed. Thus, almost every hit goes to the origin.
- ok123456 1y agoThe contents in question are statically generated, 1-3 KB HTML files. Hosting a single image would be the equivalent of cold serving 100s of requests. Putting up a scraper shield seems like it's more of a political statement than a solution to a real technical problem. It's also antithetical to open collaboration and an open internet of which Linux is a product.
- whatevaa 1y agoBots don't respect that.
- 1gn15 1y agoUse a CDN.
- trenchpilgrim 1y agoA great option for most people, and indeed Anubis' README recommends using Cloudflare if possible. However, not everyone can use a paid CDN. Some people can't pay because their payment methods aren't accepted. Some people need to serve content or to countries which a major CDN can't for legal and compliance reasons. Some organizations need their own independent infrastructure to serve their organizational misson.
- Aachen 1y agoSo that someone else pays for your bandwidth while seeing who is interested in this content? Idk about that solution
- pinoy420 1y ago[dead]
- herf 1y agoWe deployed hashcash for a while back in 2004 to implement Picasa's email relay - at the time it was a pretty good solution because all our clients were kind of similar in capability. Now I think the fastest/slowest device is a broader range (just like Tavis says), so it is harder to tune the difficulty for that.
- 0003 1y agoSoon any attempt to actually do it would indicate you're a bot.
- alt187 1y ago[flagged]
- fortran77 1y agoExactly right. Few here get it because everyone here climbs over each other to see who can virtue signal the most.
- sethaurus 1y agoCould you expand on what ideology this tool is broadcasting and what virtue is being signalled?
- haskellshill 1y ago>Please call me (order of preference): They/them or She/her please. Take a wild guess
- Deestan 1y agoPlease show me on the doll where this stranger's personal identity hurt you.
- haskellshill 1y agoThey asked a question about what ideology was being referred to, and you're angry I clarified?
- haskellshill 1y agoAlso, you do know the > Please show me on the doll where this stranger hurt you phrasing is pretty closely associated with child abuse investigations, right? I don't know why you'd associate gender identity with that?
- anonfordays 1y agoAnd these Nazis, are they in the room with us now?
- userbinator 1y agoAs I've been saying for a while now - if you want to filter for only humans, ask questions only a human can easily answer; counting the number of letters in a word seems to be a good way to filter out LLMs, for example. Yes, that can be relatively easily gotten around, just like Anubis, but with the benefit that it doesn't filter out humans and has absolutely minimal system requirements (a browser that can submit HTML forms), possibly even less than the site itself. There are forums which ask domain-specific questions as a CAPTCHA upon attempting to register an account, and as someone who has employed such a method, it is very effective. (Example: what nominal diameter is the intake valve stem on a 1954 Buick Nailhead?)
- cm2012 1y agoThere is a decent segment of the population that will gave a hard time with that.
- wavemode 1y agoSo it's no different from real CAPTCHAs, then.
- soared 1y agoTried and true method! An old video game forum named moparscape used to ask what mopar was and I always had to google it
- Aachen 1y agoGood thing modern bots can't do a web search!
- userbinator 1y agoThey will be as likely if not more so to fall victim to the large amount of misinformation... and AI-generated crap you'll find from doing so.
- ack_complete 1y ago
- spiritplumber 1y agoFor the same reason why cats sit on your keyboard. Because they can
- galaxyLogic 1y agoI think the solution to captcha-rot is micro-payments. It does consume resources to serve a web-page so whose gonna pay for that? If you want to do advertisement then don't require a payment, and be happy that crawlers will spread your ad to the users of AI-bots. If you are a non-profit-site then it's great to get a micro-payment to help you maintain and run the site.
- auggierose 1y agoWould it not be more effective just to require payment for accessing your website? Then you don't need to care about bot or not.
- eqvinox 1y agoTFA — and most comments here — seem to completely miss what I thought was the main point of Anubis: it counters the crawler's "identity scattering"/sybil'ing/parallel crawling. Any access will fall into either of the following categories: - client with JS and cookies. In this case the server now has an identity to apply rate limiting to, from the cookie. Humans should never hit it, but crawlers will be slowed down immensely or ejected. Of course the identity can be rotated — at the cost of solving the puzzle again. - amnesiac (no cookies) clients with JS. Each access is now expensive. (- no JS - no access.) The point is to prevent parallel crawling and overloading the server. Crawlers can still start an arbitrary number of parallel crawls, but each one costs to start and needs to stay below some rate limit. Previously, the server would collapse under thousands of crawler requests per second. That is what Anubis is making prohibitively expensive.
- thayne 1y agoYou don't necessarily need JS, you just need something that can detect if Anybis is used and complete the challenge.
- deleted 1y ago[deleted]
- eqvinox 1y agoSure, doesn't change anything though; you still need to spend energy on a bunch of hash calculations.
- rocqua 1y agoBut then you rate limit that challenge. You could setup a system for parellelizing the creation of these Anubis PoW cookies independent of the crawling logic. That would probably work, but it's a pretty heavy lift compared to 'just run a browser with JavaScript'.
- rocqua 1y agoThis is a good point, presuming the rate limiting is actually applied.
- thayne 1y agoI can't find any documentation that says Anubis does this, (although it seems odd to me that it wouldn't, and I'd love a reference) but it could do the following: 1. Store the nonce (or some other identifier) of each jwt it passes out in the data store 2. Track the number or rate of requests from each token in the data store 3. If a token exceeds the rate limit threshold, revoke the token (or do some other action, like tarpit requests with that token, or throttle the requests) Then if a bot solves the challenge it can only continue making requests with the token if it is well behaved and doesn't make requests too quickly. It could also do things like limit how many tokens can be given out to a single ip address at a time to prevent a single server from generating a bunch of tokens.
- ChocolateGod 1y agoI have a S24 (flagship of 2024) and Anubis often takes 10-20 seconds to complete, that time is going to add up if more and more sites adopt it, leaning to a worse browsing experience and wasted battery life. Meanwhile AI farms will just run their own nuclear reactors eventually and be unaffected. I really don't understand why someone thought this was a good idea, even if well intentioned.
- vova_hn 1y agoI have Pixel 7 (released in 2022) and it usually takes less than a second...
- whatevaa 1y agoSomething is wrong with your flagship if it takes that long.
- prmoustache 1y agoI guess his flagship IS compromised and part of an AI crawling botnet ;-)
- ChocolateGod 1y agoSamsung's UI has this feature where it turns on power saving mode when it detects light use.
- Lammy 1y agoYou're looking at it wrong.
- prmoustache 1y agoSomething must be wrong on your flagship smartphone because I have an entry level one that doesn't take that long. It seems there is a large number of operations crawling the web to build models that aren't using directly infrastructure hosted on AI farms BUT botnet running on commodity hardware and residencial networks to circumvent their ip range from being blacklisted. Anubis point is to block those.
- 1y ago
- a-dub 1y agoblame canada
- whatevaa 1y agoSite doesn't load, must be hit by AI crawlers.
- a-dub 1y agoi suppose one nice property is that it is trivially scalable. if the problem gets really bad and the scrapers have llms embedded in them to solve captchas, the difficulty could be cranked up and the lifetime could be cranked down. it would make the user experience pretty crappy (party like it's 1999) but it could keep sites up for unauthenticated users without engaging in some captcha complexity race. it does have arty political vibes though, the distributed and decentralized open source internet with guardian catgirls vs. late stage capitalism's quixotic quest to eat itself to death trying to build an intellectual and economic robot black hole.
- deevus 1y agoThis seems like a good place to ask. How do I stop bots from signing up to my email list on my website without hosting a backend?
- account42 1y agoDepending on your target audience you could require people signing up to send you and email first.
- Tractor8626 1y ago[flagged]
- nicman23 1y agodid you read the article?
- voidUpdate 1y agoIf it didn't work, do you think it would be so widespread?
- anonfordays 1y agoJust use Anubis Bypass: https://addons.mozilla.org/en-US/android/addon/anubis-bypass/ https://addons.mozilla.org/en-US/android/addon/anubis-bypass... Haven't seen dumb anime characters since.
- miohtama 1y agoThe solution is to make premium subscription service for those who do not want to solve CAPTCHAs. Money is the best proof of humanity.
- lock1 1y agoIsn't that line of reasoning implies companies with multi-billion dollars in their war chest are much more "human" than a literal human with student loans?
- PeterStuer 1y ago[flagged]
- Biganon 1y ago[flagged]
- PeterStuer 1y agoWhat's the point in having Karma if you never use it?
- deleted 1y ago[deleted]
- zaptrem 1y agoIf people are truly concerned about the crawlers hammering their 128mb raspberry pi website then a better solution would be to provide an alternative way for scrapers to access the data (e.g., voluntarily contribute a copy of their public site to something like common crawl). If Anubis blocked crawler requests but helpfully redirected to a giant tar ball of every site using their service (with deltas or something to reduce bandwidth) I bet nobody would bother actually spending the time to automate cracking it since it’s basically negative value. You could even make it a torrent so most of the be costs are paid by random large labs/universities. I think the real reason most are so obsessed with blocking crawlers is they want “their cut”… an imagined huge check from OpenAI for their fan fiction/technical reports/whatever.
- elsjaako 1y agoThere's a lot of people that really don't like AI, and simply don't want their data used for it.
- zaptrem 1y agoWhile that’s a reasonable opinion to have, it’s a fight they can’t really win. It’s like putting up a poster in a public square then running up to random people and shouting “no, this poster isn’t for you because I don’t like you, no looking!” Except the person they’re blocking is an unstoppable mega corporation that’s not even morally in the wrong imo (except for when they overburden people’s sites, that’s bad ofc)
- guappa 1y agoThe looking is fine, the photographing and selling the photo less so… and fyi in denmark monuments have copyright so if you photograph and sell the photos you owe fees :)
- sussmannbaka 1y agoNo, this doesn’t work. Many of the affected sites have these but they’re ignored. We’re talking about git forges, arguably the most standardised tool in the industry, where instead of just fetching the repository every single history revision of every single file gets recursively hammered to death. The people spending the VC cash to make the internet unusable right now don’t know how to program. They especially don’t give a shit about being respectful. They just hammer all the sites, all the time, forever.
- pkal 1y agoSuperficial comment regarding the catgirl, I don't get why some people are so adamant and enthusiastic for others to see it, but if you like me find it distasteful and annoying, consider copying these uBlock rules: https://sdf.org/~pkal/src+etc/anubis-ublock.txt https://sdf.org/~pkal/src+etc/anubis-ublock.txt. Brings me joy to know what I am not seeing whenever I get stopped by this page :)
- squigz 1y agoI don't get why so many people find it "distasteful and annoying"
- account42 1y agoYou could respect it without "getting" it though.
- IshKebab 1y agoI can't really explain it but it definitely feels extremely cringeworthy. Maybe it's the neckbeard sexuality or the weird furry aspect. I don't like it.
- pkal 1y agoCan you clarify if you mean that you do no understand the reasons that people dislike these images, or do you find the very idea of disliking it hard to relate to? I cannot claim that I understand it well, but my best guess is that these are images that represent a kind of culture that I have encountered both in real-life and online that I never felt comfortable around. It doesn't seem unreasonable that this uneasiness around people with identity-constituting interests in anime, Furries, MLP, medieval LARP, etc. transfers back onto their imagery. And to be clear, it is not like I inherently hate anime as a medium or the idea of anthropomorphism in art. There is some kind of social ineptitude around propagating these _kinds_ of interests that bugs me. I cannot claim that I am satisfies with this explanation. I know that the dislike I feel for this is very similar to that I feel when visiting a hacker space where I don't know anyone. But I hope that I could at least give a feeling for why some people don't like seeing catgirls every time I open a repository and that it doesn't necessarily have anything to do with advocating for a "corporate soulless web".
- anarki8 1y agoArticle might be a bit shallow, or maybe my understanding of how Anubis works is incorrect? 1. Anubis makes you calculate a challenge. 2. You get a "token" that you can use for a week to access the website. 3. (I don't see this being considered in the article) "token" that is used too much is rate limited. Calculating a new token for each request is expensive.
- jeroenhd 1y agoThat's the basic principle. It's a tool to fight to crawlers that spam requests without cookies to prevent rate limiting. The Chinese crawlers seem to have adjusted their crawling techniques to give their browsers enough compute to pass standard Anubis checks.
- Aachen 1y agoThat, but apparently also restrictions on what tech you can use to access the website: - https://news.ycombinator.com/item?id=44971990 https://news.ycombinator.com/item?id=44971990 person being blocked with `message looking something like "you failed"` - https://news.ycombinator.com/item?id=44970290 https://news.ycombinator.com/item?id=44970290 mentions of other requirements that are allegedly on purpose to block older clients (as browser emulators presumably often would appear to be, because why would they bother implementing newer mechanisms when the web has backwards compatibility)
- account42 1y agoGood on you that you found a solution to myself but personally I will just not use websites that pull this and not contribute to projects where using such a website is required. If you respect me so little that you will make demands about how I use my computer and block me as a bot if I don't comply then I am going to assume that you're not worth my time.
- russelg 1y agoInteresting take to say the Linux Kernel is not worth your time.
- account42 1y agoAs far as I know Linux kernel contributions still use email.
- anarki8 1y agoThis sounds a bit overdramatic for a less than a second waiting time per week for each device. Unless you employ an army of crawlers of course.
- KolmogorovComp 1y agoWhy does Anubis not leverage PoW from its users to do something useful (at best, distributed computing for science, at worst, a crypto-currency at least allowing the webmasters to get back some cash)
- johnklos 1y agoPeople are already complaining. Could you imagine how much fodder this'd give people who didn't like the work or the distribution of any funds that a cryptocurrency would create (which would be pennies, I think, and more work to distribute than would be worth doing).
- est 1y agoI hope there's some kind of memory-hungry checker to replace the CPU cost. a 2GB memory consumption wont stop them, but it will limit the parallelism of crawlers.
- pembrook 1y agoSomething feels bizarrely incongruent about the people using Anubis. These people used to be the most vehemently pro-piracy, pro internet freedom and information accessibility, etc. Yet now when it's AI accessing their own content, suddenly they become the DMCA and want to put up walls everywhere. I'm not part of the AI doomer cult like many here, but it would seem to me that if you publish your content publicly, typically the point is that it would be publicly available and accessible to the world...or am I crazy? As everything moves to AI-first, this just means nobody will ever find your content and it will not be part of the collective human knowledge. At which point, what's the point of publishing it.
- SnuffBox 1y agoIt is rather funny. "We must prevent AI accessing the Arch Linux help files or it will start the singularity and kill us all!"
- GreenWatermelon 1y agoIn case you're genuinely confused, the reason for Anubis and similar tools is that AI-training-data-scraping crawlers are assholes, and strangle the living shit out of any webserver they touch, like a cloud of starving locusts descending upon a wheat field. i.e. it's DDoS protection.
- walthamstow 1y ago> I host this blog on a single core 128MB VPS Where does one even find a VPS with such small memory today?
- tambourine_man 1y agoOr software to run on it. I'm intrigued about this claim as well.
- Aachen 1y agoThe software is easy. Apt install debian apache2 php certbot and you're pretty much set to deploy content to /var/www. I'm sure any BSD variant is also fine, or lots of other software distributions that don't require a graphical environment On an old laptop running Windows XP (yes, with GUI, breaking my own rule there) I've also run a lot of services, iirc on 256MB RAM. XP needed about 70 I think, or 52 if I killed stuff like Explorer and unnecessary services, and the remainder was sufficient to run a uTorrent server, XAMPP (Apache, MySQL, Perl and PHP) stack, Filezilla FTP server, OpenArena game server, LogMeIn for management, some network traffic monitoring tool, and probably more things I'm forgetting. This ran probably until like 2014 and I'm pretty sure the site has been on the HN homepage with a blog post about IPv6. The only thing that I wanted to run but couldn't was a Minecraft server that a friend had requested. You can do a heck of a lot with a hundred megabytes of free RAM but not run most Javaware :)
- tambourine_man 1y agoWhat I meant is that I’m not sure it will even boot. Bookworm minimum requirements are 256MB of RAM. https://www.debian.org/releases/bookworm/armel/ch03s04.en.html?utm_source=chatgpt.com https://www.debian.org/releases/bookworm/armel/ch03s04.en.ht... 128MB should be plenty. I used systems for years with much less. But in reality, Linux is much heavier these days.
- grahar64 1y agoI write about something similar a while back https://maori.geek.nz/proof-of-human-2ee5b9a3fa28 https://maori.geek.nz/proof-of-human-2ee5b9a3fa28 About the difficulty of proving you are human especially when every test built has so much incentive to be broken. I don't think it will be solved, or could ever be solved.
- zoobab 1y agoTime to switch to stagit. Unfortunately it does not generate static pages for a git repo except "master". I am sure someone will modify to support branches.
- wraptile 1y agoI'm a scraper developer and Anubis would have worked 10 - 20 years ago, but now all broad scrapers run on a real headless browser with full cookie support and costs relatively nothing in compute. I'd be surprised if LLM bots would use anything else given the fact that they have all of this compute and engineers already available. That being said, one point is very correct here - by far the best effort to resist broad crawlers is a _custom_ anti-bot that could be as simple as "click your mouse 3 times" because handling something custom is very difficult in broad scale. It took the author just few minutes to solve this but for someone like Perplexity it would take hours of engineering and maintenance to implement a solution for each custom implementation which is likely just not worth it. You can actually see this in real life if you google web scraping services and which targets they claim to bypass - all of them bypass generic anti-bots like Cloudflare, Akamai etc. but struggle with custom and rare stuff like Chinese websites or small forums because scraping market is a market like any other and high value problems are solved first. So becoming a low value problem is a very easy way to avoid confrontation.
- hahn-kev 1y agoBot blocking through obscurity
- lbhdc 1y agoThat's really the only option available here, right? The goal is to keep sites low friction for end users while stopping bots. Requiring an account with some moderation would stop the majority of bots, but it would add a lot of friction for your human users.
- brookst 1y agoThe other option is proof of work. Make clients use JS to do expensive calculations that aren’t a big deal for single clients, but get expensive at scale. Not ideal, but another tool to potentially use.
- tovej 1y agoI like it, make the bot developers play whack-a-mole. Of course, you're going to have to verify each custom puzzle aren't you.
- Aachen 1y ago> an AI vendor will have a datacenter full of compute capacity. It feels like this solution has the problem backwards, effectively only limiting access to those without resources Sure, if you ignore that humans click on one page and the problematic scrapers (not the normal search engine volume, but the level we see nowadays where misconfigured crawlers go insane on your site) are requesting many thousands to millions of times more pages per minute. So they'll need many many times the compute to continue hammering your site whereas a normal user can muster to load that one page from the search results that they were interested in
- tortillasauce 1y agoAnubis works because AI crawlers do very little requests from an ip address to bypass rate-limiting. Last year they could still be blocked by ip range, but now the requests are from so many different networks that doesn't work anymore. Doing the proof-of-work for every request is apparently too much work for them. Crawlers using a single ip, or multiple ips from a single range are easily identifiable and rate-limited.
- pluc 1y agoCan we talk about the "sexy anime girl" thing? Seems it's popular in geek/nerd/hacker circles and I for one don't get it. Browsing reddit anonymously you're flooded with near-pornographic fan-made renders of these things, I really don't get the appeal. Can someone enlighten me?
- andai 1y agoProbably depends on the person, but this stuff is mostly the cute instinct, same as videos of kittens. "Aww" and "I must protect it."
- abustamam 1y agoIt's a good question. Anime (like many media, but especially anime) is known to have gratuitous fan service where girls/women of all ages are in revealing clothing for seemingly no reason except to just entice viewers. The reasoning is that because they aren't real people, it's okay to draw and view images of anime, regardless of their age. And because geek/nerd circles tend not to socialize with real women, we get this over-proliferation of anime girls.
- pluc 1y agoThis also was my best guess. A "victimless crime" kind of logic that really really creeps me out.
- abustamam 1y agoIt is a bit unsettling, but at the risk of false dichotomy, I'd rather them get off on cartoon girls or women than getting their fill with real underage girls or otherwise unconsenting women.
- dominick-cc 1y ago2D girls don't nag and I've never had to clear their clogged hair out of my shower drain.
- SnuffBox 1y ago
- m-p-3 1y agoAnd Codeberg, even behind Anubis, is not immune from scrapers either https://social.anoxinon.de/@Codeberg/115033782514845941 https://social.anoxinon.de/@Codeberg/115033782514845941
- ajsnigrutin 1y agoI always wondered about these anti bot precautions... as a firefox user, with ad blocking and 3rd party cookies disabled, i get the goddamn captcha or other random check (like this) on a bunch of pages now, every time i visit them... Is it worth it? Millions of users wasting cpu and power for what? Saving a few cents on hosting? Just rate limit requests per second per IP and be done. Sooner or later bots will be better at captchas than humans, what then? What's so bad with bots reading your blog? When bots evolve, what then? UK style, scan your ID card before you can visit? The internet became a pain to use... back in the time, you opened the website and saw the content. Now you open it, get an antibot check, click, forward to the actual site, a cookie prompt, multiple clicks, then a headline + ads, scroll down a milimeter... do you want to subscribe to a newsletter? Why, i didn't even read the first sentence of the article yet... scroll down.. chat with AI bot popup... a bit further down login here to see full article... Most of the modern web is unusable. I know I'm ranting, but this is just one of the pieces of a puzzle that makes basic browsing a pain these days.
- trostaft 1y agoI actually really liked seeing the mascot. Brought a sense of whimsy to the Internet that I've missed for a long time.
- verall 1y agoIt's posts like this that make me really miss the webshit weekly
- Wowhappyfun 1y ago[flagged]
- usbpoet 1y agoI don't think I've ever actually seen Anubis once. Always interesting to see what's going on in parts of the internet you aren't frequenting.
- dominick-cc 1y agoI read hackernews on my phone when I'm bored and I've seen it a lot lately. I don't think I've ever seen it on my desktop.
- SnuffBox 1y agoWhenever I see an otherwise civil and mature project utilize something outwardly childish like this I audibly groan and close the page. I'm sure the software behind it is fine but the imagery and style of it (and the confidence to feature it) makes me doubt the mental credibility/social maturity of anybody willing to make it the first thing you see when accessing a webpage. Edit: From a quick check of the "CEO" of the company, I was unsurprised to have my concerns confirmed. I may be behind the times but I think there are far too many people in who act obnoxiously (as part of what can only be described as a new subculture) in open source software today and I wish there were better terms to describe it.
- buyucu 1y agoWe're 1-2 years away from putting the entire internet behind Cloudflare, and Anubis is what upsets you? I really don't get these people. Seeing an anime catgirl for 1-2 seconds won't kill you. It might save the internet though. The principle behind Anubis is very simple: it forces every visitor to brute force a math problem. This cost is negligible if you're running it on your computer or phone. However, if you are running thousands of crawlers in parallel, the cost adds up. Anubis basically makes it expensive to crawl the internet. It's not perfect, but much much better than putting everything behind Cloudflare.
- 8cvor6j844qw_d6 1y agoOn my daily browser with V8 JIT disabled, Cloudflare Turnstile has the worst performance hit, and often requires an additional click to clear. Anubis usually clears in with no clicks and no noticeable slowdown, even with JIT off. Among the common CAPTCHA solutions it's the least annoying for me.