24 ms·
Anubis saved our websites from a DDoS attack
- ranger_danger 1y agoSeems like rate-limiting expensive pages would be much easier and less invasive. Also caching... And I would argue Anubis does nothing to stop real DDoS attacks that just indiscriminately blast sites with tens of gbps of traffic at once from many different IPs.
- bastawhiz 1y agoRate limiting does nothing when your adversary has hundreds or even thousands of IPs. It's trivial to pay for residential proxies.
- supportengineer 1y agoWhy aren't there any authorities going after this problem?
- eikenberry 1y agoThey could be doing it legally.
- deleted 1y ago[deleted]
- danielheath 1y agoMost of the "free" analytics tools for android/iOS are "funded" by running residential / "real user" proxies. They wait until your phone is on wifi / battery, then make requests on behalf of whoever has paid the analytics firm for access to 'their' residential IP pool.
- marginalia_nu 1y agoThese residential botnets are pretty difficult to shut down, and often operated out of countries with poor diplomatic relations with the west.
- GoblinSlayer 1y agohttps://infatica-sdk.io/ https://infatica-sdk.io/ INFATICA LTD Reg. No.: 14863491 Unit A, 82 James Carter Road, Mildenhall, Suffolk, IP28 7DE, United Kingdom
- o11c 1y agoBecause in a "free" nation, that means "free to run malware" not "free from malware". By far most malware is legal and a portion of its income is used to fund election campaigns.
- okanat 1y ago1. Which authorities? 2. The US is currently broken and they are not going to punish only, albeit unsustainable, growth in their economy. 3. Internet is global. Even EU wants to regulate, will they charge big tech leaders and companies with information tech crimes which will pierce the corporate veil? It will ensure that nobody will invest in unsustainable AI growth in the EU. However fucking up economy and the planet is how the world operates now, and without infinite growth you lose buying power for everything. So everybody else will continue to do fuckery. 4. What can a regulating body do? Force disconnects for large swaths of internet? Then Internet is no more.
- folmar 1y agoI would go for making the AI companies pay. Identifying end users for other abuse works but there are problems on state borders, for monetary solutions it should be easier.
- Ocha 1y agoRate limit according to what? It was 35k residential IPs. Rate limit would end up keeping real users out.
- linsomniac 1y agoRate limit according to destination URL (the expensive ones), not source IP. If you have expensive URLs that you can't serve more than, say 3 of at a time, or 100 of per minute, NOT rate limiting them will end up keeping real users out simply because of the lack of resources.
- pluto_modadic 1y agothis feels like something /you can do on your servers/, and that other folks with resource constraints (like time, budget, or the hardware they have) find anubis valuable.
- linsomniac 1y agoSure, didn't mean to imply Anubis wasn't an alternative, just was clearing up that there are options beyond source IP rate limiting, which several people seemed to be thinking was the only option because of comments about rate limiting not working because it was coming from 35K IP addresses.
- danielheath 1y agoRight - but if you have, say, 1000 real user requests for those endpoints daily, and thirty million bot requests for those endpoints, the practical upshot of this approach is that none of the real users get to access that endpoint.
- Groxx 1y agoYeah, at that point to might as well just turn off the servers. It's even cheaper at cutting off requests, and it'll serve just as many legitimate users.
- PaulDavisThe1st 1y agoIn the last two months, ardour.org's instance of fail2ban has blocked more than 1.2M distinct IP addresses that were trawling our git repo using http instead of just fetching the goddam repository. We shut down the website/http frontend to our git repo. There are still 20k distinct IP addresses per day hitting up a site that issues NOTHING but 404 errors.
- lousken 1y agoYes, have everything static (if you can't, use caching), optimize images, rate limit anything you have to generate dynamically
- felsqualle 1y agoHi, author here. Caching is already enabled, but this doesn’t work for the highly dynamic parts of the site like version history and looking for recent changes. And yes, it doesn’t work for volumetric attacks with tens of gbps. At this point I don’t think it is a targeted attack, probably a crawler gone really wild. But for this pattern, it simply works.
- GoblinSlayer 1y agoThere's a theory they didn't get through, because it's a new protection method and the bots don't run javascript. It could be as simple as <script>setCookie("letmein=1");reload();</script>
- toast0 1y ago> And I would argue Anubis does nothing to stop real DDoS attacks that just indiscriminately blast sites with tens of gbps of traffic at once from many different IPs. Volumetric DDoS and application layer DDoS are both real, but volumetric DDoS doesn't have an opportunity for cute pictures. You really just need a big enough inbound connection and then typically drop inbound UDP and/or IP fragments and turn off http/3. If you're lucky, you can convince your upstream to filter out UDP for you, which gives you more effective bandwidth.
- herpdyderp 1y agoCan Anubis be restyled to be more... professional? I like the playfulness, but I know at least some of my clients will not.
- ranger_danger 1y agoyes it's open source https://git.kernel.org/ https://git.kernel.org/ changed theirs
- LPisGood 1y agoI’ve heard people say that before. They would love to use it if there wasn’t a playful animated character. The code is open source, so I can’t imagine making a fork to remove that is a Herculean effort.
- unsnap_biceps 1y agoWhen I last looked into it, they are planning a white label service to customize the look and has been requesting folks to not fork and modify the images. > Regardless, Xe did ask nicely to not change out the images shipped as a whitelabel service is planned in the future https://github.com/TecharoHQ/anubis/pull/204#issuecomment-2776558289 https://github.com/TecharoHQ/anubis/pull/204#issuecomment-27...
- natebc 1y agoIt's also mentioned on the docs site: https://anubis.techaro.lol/docs/funding/ https://anubis.techaro.lol/docs/funding/
- LPisGood 1y agoThat’s the beautiful thing about open source, they ask but do not demand. Of course, if you use this service for your enterprise, the Right Thing To Do would be support the excellent project financially, but this is by no means required. If you want to use this project on your site and don’t like the logo, you are free to change it. If the site is personal and this project is not something you would spend money on, I don’t even think it is unethical to change the image.
- chrisnight 1y ago> Solving the challenge–which is valid for one week once passed– One thing that I've noticed recently with the Arch Wiki adding Anubis, is that this one week period doesn't magically fix user annoyances with Anubis. I use Temporary Containers for every tab, which means that I constantly get Anubis regenerating tokens, since the cookie gets deleted as soon as the tab is closed. Perhaps this is my own problem, but given the state of tracking on the internet, I do not feel it is an extremely out-of-the-ordinary circumstance to avoid saving cookies.
- jsheard 1y agoIt could be worse, the main alternative is something like Cloudflares death-by-a-thousand-CAPTCHAs when your browser settings or IP address put you on the wrong side of their bot detection heuristics. Anubis at least doesn't require any interaction to pass. Unfortunately nobody has a good answer for how to deal with abusive users without catching well behaved but deliberately anonymous users in the crossfire, so it's just about finding the least bad solution for them.
- lousken 1y agoI hated everyone who enabled the cloudflare validation thing on their website, because it was blocked for months (I got stuck on that captcha that was refusing my Firefox). Eventually they fixed it but it was really annoying.
- goku12 1y agoThe CF verification page still appears far too often in some geographic regions. It's such an irritant that I just close the tab and leave when I see it. It's so bad that seeing the Anubis page instead is actually a big relief! I consider the CF verification and its enablers as a shameless attack the open web - a solution nearly as bad as the problem it tries to solve.
- _bin_ 1y agoForget esoteric areas, I'm an average American guy who gets them running from a residential IP or cell IP. It even happens semi-frequently on my iPhone which is insane. I guess I must have "bot-like" behavior in my browsing, even from a cell.
- Tiberium 1y agoFrom looking at some of the rules like https://github.com/TecharoHQ/anubis/blob/main/data/bots/headless-browsers.yaml https://github.com/TecharoHQ/anubis/blob/main/data/bots/head... it seems that Anubis explicitly punishes bots that are "honest" about their user agent - I might be missing something, but isn't this just pressuring anyone who does anything bot-related to just lie about their user agent? Flat out user-agent blacklist seems really weird, it's going to reward the companies that are more unethical in their scraping practices than the ones who report their user agent truthfully. From the repo it also seems like all the AI crawlers are also DENY, which, again, would reward AI companies that don't disclose their identity in the user agent.
- userbinator 1y agoUser-agent header is basically useless at this point. It's trivial to set it to whatever you want, and all it does is help the browser incumbents.
- Tiberium 1y agoYou're right, that's why I'm questioning the reason Anubis implemented it this way. Lots of big AI companies are at least honest about their crawlers and have proper user agents (which Anubis outright blocks). So "unethical" companies who change the user-agent to something normal will have an advantage with the way Anubis is currently set up by default. I'm aware that end users can modify the rules, but in reality most will just use the defaults.
- xena 1y agoShitty heuristics buy time to gather data and make better heuristics.
- MillironX 1y agoDespite broadcasting their user agents properly, the AI companies ignore robots.txt and still waste my server resources. So yeah, the dishonest botnets will have an advantage, but I don't give swindlers a pass just because they rob me to my face. I'm okay with defaults that punish all bots.
- justusthane 1y agoI don’t really understand why this solved this particular problem. The post says: > As an attacker with stupid bots, you’ll never get through. As an attacker with clever bots, you’ll end up exhausting your own resources. But the attack was clearly from a botnet, so the attacker isn’t paying for the resources consumed. Why don’t the zombie machines just spend the extra couple seconds to solve the PoW (at which point, they would apparently be exempt for a week and would be able to continue the attack)? Is it just that these particular bots were too dumb?
- judge2020 1y agoAnubis is new, so there may not have been foresight to implement a solver to get around it. Also, I wouldn't be surprised if the botnet actor is using vended software, not making it themselves to where they could quickly implement a solver to continue their attack.
- cbarrick 1y agoI think the explanation "you’ll end up exhausting your own resources" is wrong for this case. I think you are correct that the bots are simply too dumb. The likely explanation is that the bots are just curling the expensive URLs without a proper JavaScript engine to solve the challenge. E.g. if I hack a bunch of routers around the world to act as my botnet, I probably wouldn't have enough storage to install Chrome or Selenium. The lightweight solution is just to use curl/wget (which may be pre-installed) or netcat/telnet.
- maeln 1y agoMost DDoS bot don't bother running JS. A lot of botnets don't even really allow it, because the malware they run on the infected target only allow for basic stuff like simple HTTP request. This is why they often do some reconnaissance to find pages that take a long time to load, and therefore are probably using a lot of I/O and/or CPU time on the target server. Then they just spam the request. Huge botnet don't even bother with all that, they just kill you with the bandwidth.
- tpool 1y agoIt's so bad we're going to the old gods for help now. :)
- Hamuko 1y agoI’d sic Yogg-Saron on these scrapers if I could.
- rubyn00bie 1y agoSort of tangential but I’m surprised folks are still using Apache all these years later. Is there a certain language that makes it better than Nginx? Or it just the ease of use configuration that still pulls people? I switched to Nginx I don’t even know how many years ago and never looked back, just more or less wondering if I should.
- anotherevan 1y agoEqually tangential, but I switched form Nginx to Caddy a few years ago and never looked back.
- ahofmann 1y agoI'm using nginx since what feels like decades and occasionally I miss the ability to use .htaccess files. This is a very nice way to configure stuff on a server.
- felsqualle 1y agoI use it because that’s the one I’m most familiar with. Using it since 15 years and counting. And since it doesn’t the job for me, I never had the urge to look into alternatives.
- mrweasel 1y agoApache does everything, it's fairly easy to configure. If there's something you want to do, Apache mostly knows how, or have a module. If you run a fleet of servers, all doing different things, Apache is a good choice because all the various uses are going to be supported. It might not be the best choice in each individual case, but it is the one that works in all of them. I don't know why some are so quick to write off Apache. Is just because it's old? It's still something like the second most used webserver in the world.
- forinti 1y agoApache has so much functionality. Why wouldn't anybody use it? I started using it when Oracle's Webcache wouldn't support newer certificates and I had to keep Oracle Portal running. I could edit the incoming certificate (I had to snip the header and the footer) and put it in a specific header for Portal to accept it.
- gitroom 1y agoKinda love how deep this gets into the whole social contract side of open source. Honestly, it's been a pain figuring out what feels right when folks mix legal rules and personal asks.
- lytedev 1y agoYeah I had no idea that some folks would get so passionate about making changes to a piece of FOSS based on a request on a certain footer-esque documentation page. I think its a great discussion though that gets to the heart of open source and software freedom and how that can seem orthogonal to business needs depending on how you squint.
- CaptainFever 1y ago> To me, Anubis is not only a blocker for AI scrapers. Anubis is a DDoS protection. Anubis is DDoS protection, just with updated marketing. These tools have existed forever, such as CloudFlare Challenges, or https://github.com/RuiSiang/PoW-Shield https://github.com/RuiSiang/PoW-Shield. Or HashCash. I keep saying that Anubis really has nothing much to do with AI (e.g. some people might mistakenly think that it magically "blocks AI scrapers"; it only slows down abusive-rate visitors). It really only deals with DoS and DDoS. I don't understand why people are using Anubis instead of all the other tools that already exist. Is it just marketing? Saying the right thing at the right time?
- consp 1y agoKnowing something exists is half the challenge. Never used it but ,maybe ease of use/setup or license?
- immibis 1y agomarketing plus a product that Just Does The Thing, it seems like. No bullshit. btw it only works on AI scrapers because they're DDoSes.
- CaptainFever 1y agoNot all DDoSes are AI-related, and not all AI scrapers are DDoSes.
- superkuh 1y agoBut almost all DoS's we're talking about are from corporations. The real non-human danger.
- Imustaskforhelp 1y agoI agree with you that it is infact a DDOS protection but still, the fact that it is open source and created by a really cool dev (she is awesome), I think I don't really mind it gaining popularity. And also they had created it out of their own necessity which is also really nice. Anubis is getting real love out there and I think I am all for it. I personally host a lot of my stuff on cloudflare due to it being free with cloudflare workers but if I ever have a vps, I am probably going to use anubis as well
- mrweasel 1y agoSadly it hard to tell if this is an actual DDoS attack, or scrappers descending on the site. It all looks very similar. The search engines always seemed happy to announce that they are in fact GoogleBot/BingBot/Yahoo/whatever and frequently provided you with their expected IP ranges. The modern companies, mostly AI companies, seems to be more interested in flying under the radar, and have less respect for the internet infrastructure at a whole. So we're now at a point where I can't tell if it's an ill willed DDoS attack or just shitty AI startup number 7 reloading training data.
- piokoch 1y agoYes, search engines were not hiding, as it was website owner interest involved here as well - without those search bots their sites would not be indexed and searchable in the Internet. So there was kind of win-win situation, in most typical cases at least, as for instance publishers complained about deep links, etc. because their ads revenue was hurt. AI scrapping bots provide zero value for sites owners.
- Valodim 1y agoIs this really true? If I have a marketing website for a product, isn't it in my interest to have that marketing incorporated in AI models?
- jeroenhd 1y ago> The modern companies, mostly AI companies, seems to be more interested in flying under the radar, and have less respect for the internet infrastructure at a whole I think that makes a lot of sense. Google's goal is (or perhaps used to be) providing a network of links. The more they scrape you, the more visitors you may end up receiving, and the better your website performs (monetarily, or just in terms of providing information to the world). With AI companies, the goal is to consume and replace. In their best case scenario, your website will never receive a visitor again. You won't get anything in return for providing content to AI companies. That means there's no reason for website administrators to permit the good ones, especially for people who use subscriptions or ads to support their website operating costs.
- anonfordays 1y agoLooks similar to haproxy-protection: https://gitgud.io/fatchan/haproxy-protection/ https://gitgud.io/fatchan/haproxy-protection/
- fatchan 1y agoHey, funny to see my project mentioned here also. Yes, similar in concept. Some differences: - Uses HAProxy (duh) - Proof of work can be either sha256 or argon2 - Optional recaptcha/hcaptcha in addition to the proof of work - Includes a script for your page that will re-solve the challenge in the background before the cookie expires There's also a control panel, dns server, etc. I kinda built my own everything because I refused to use bunny/cloudflare/whatever. One thing I will say though, is that proof-of-work alone isn't a solution for ddos mitigation and bot protection! I've seen attackers using a mass of proxies and headless browsers to solve the challenge, or even writing code to extract and solve the challenge directly (https://github.com/lizthegrey/tor-fetcher https://github.com/lizthegrey/tor-fetcher). To adequately protect against more targeted attacks, you need additional acl and heuristics, browser fingerprinting, tls fingerprinting, ip reputation, etc. I do offer the whole thing setup as a commercial service, but will refrain from too much shilling. It's fun, and I love seeing similar softwares help fight the horde of AI scrapers :^)
- anonfordays 1y ago>One thing I will say though, is that proof-of-work alone isn't a solution for ddos mitigation and bot protection! I've seen attackers using a mass of proxies and headless browsers to solve the challenge If you make the challenge sufficiently difficult enough, it should mitigate this no? >or even writing code to extract and solve the challenge directly (https://github.com/lizthegrey/tor-fetcher https://github.com/lizthegrey/tor-fetcher). Similarly if the challenge is difficult, wouldn't matter where it's solved. I'm not sure why one would use Anubis over haproxy-protection.
- forty 1y agoAnubis is nice, but could we have a PoW system integrated in protocols (http or TLS, I'm not sure) so we don't have to require JS ?
- fc417fc802 1y agoProtocol is the wrong level. Integrate with the browser. Add a PoW challenge header to the HTTP response, receive a POW solution header with the next request.
- forty 1y agoI think you've just described a protocol ;) Yes it could be in higher layer than what I suggested indeed, on top of HTTP sounds good to me. My rule of thumb is that it should work with curl (which makes it not antibots, but just anti scrapper & ddos, which is what I have a problem with)
- fc417fc802 1y agoAh yeah sloppy wording on my part. I think it should ideally be its own protocol built on top as opposed to integrated into an existing one. Integration is good but mandatory complexity and tight coupling not so much.
- selfhoster11 1y agoI'd much prefer for this to be standardised rather than an ad-hoc layer on top of what we have. Our protocols are already complex, and at least what we would be doing is moving that complexity somewhere where it can be handled more conveniently.
- fc417fc802 1y agoIt would still be standardized. Anyone who wanted to support it would. Those who didn't want to support it wouldn't be burdened. And it could then evolve on its own, gaining variants for layering it on additional underlying protocols. It's basic separation of responsibilities. It's helpful for reuse but also innovation. For example, the auth scheme baked in to HTTP is pretty much stuck in time and not very useful. We'd likely be better off if it wasn't tightly coupled to something unrelated like that. If I were implementing an HTTP stack I'd want to omit it, but that would make me noncompliant.
- vachina 1y agoIt’s not Anubis that saved your website, literally any sort of Captcha, or some dumb modal with a button to click into the real contents would’ve worked. These crawlers are designed to work on 99% of hosts, if you tweak your site just so slightly out of spec, these bots wouldn’t know what to do.
- deleted 1y ago[deleted]
- boreq 1y agoSo what you are saying is that it's anubis that saved their website.
- butz 1y agoAs usual, there is a negative side to such protection: I was trying to download some raw files from git repository and instead of data got bunch of html. After quick look it turned out to be Anubis HTML page. Another issue was with broken links to issue tickets on main page, where Anubis was asking wrapper script to solve some hashes. Lesson here: after deploying Anubis, please carefully check the impact. There might be some unexpected issues.
- eadmund 1y ago> I was trying to download some raw files from git repository and instead of data got bunch of html. After quick look it turned out to be Anubis HTML page. Yup. Anubis breaks the web. And it requires JavaScript, which also breaks the web. It’s a disaster.
- lytedev 1y agoI'm using a nearly default configuration which seems to not have this problem. curl still works and so do downloads. I guess if your cookie expired at just the right time that could cause this issue, and that might be worth thinking about, but I think "breaks the web" is overstating it a bit, at least for the default configuration.
- ziddoap 1y agoI feel like it's much more reasonable to blame the companies & people that are making it a necessity to have some sort of protection like Anubis for ruining the web (over-aggressive scrapers, bot farms, etc.), rather than blaming Anubis.
- qiu3344 1y agoAs someone who has a lot of experience with (not AI related) web scraping, fingerprinting and WAFs, I really like what Anubis is doing. Amazon, Akamai, Kasada and other big players in the WAF/Antibot industry will charge you millions for the illusion of protection and half-baked javascript fingerprint collectors. They usually calculate how "legit" your request is based on ambiguous factors, like the vendor name of your GPU (good luck buying flight tickets in a VM) or how anti-aliasing is implemented on you fonts/canvas. Total bullshit. Most web scrapers know how to bypass it. Especially the malicious ones. But the biggest reason why I'm against these kind of systems is how they support the browser mono-culture. Your UA is from Servo or Ladybird? You're out of luck. That's why the idea choosing a purely browser-agnostic way of "weighting the soul" of a request resonates highly with me. Keep up the good work!
- xena 1y agoThanks! I'm going out of my way to make sure smaller browsers like Pale Moon aren't locked out when I add reputation into the equation. One of my prototypes that would work in concert with other changes works in links too :)
- parrit 1y agoIf I see a cute cartoon with a cryptocurrency mining like "KHash/s" thing I am gonna leave that site real quick! It should explain it isn't mining and just verifying the browser or such.
- lytedev 1y agoIt includes links with explanations, but the page does kind of "fly by" in many cases. At which point, would you still leave? I'm guessing folks have seen enough captcha and CloudFlare verification pages to get a sense that they're being "soul" checked and that it's not an issue usability-wise.
- KronisLV 1y ago> We use a stack consisting of Apache2, PHP-FPM, and MariaDB to host the web applications. Oh hey, that’s a pretty utilitarian stack and I’m happy to see MariaDB be used out there. Anubis is also really cool, I do imagine that proof of work might become more prevalent in the future to deal with the sheer amount of bots and bad actors (shame that they exist) out there, albeit in the case of hijacked devices it might just slow them down, hopefully to a manageable degree, instead of IP banning them altogether. I do wonder if we’ll ever see HTTP only versions of PoW too, not just JS based options, though that might need to be a web standard or something.
- pmlnr 1y agoAnyone knows a solution that works without js?
- ximm 1y agoClient must provide a proof-of-work. There is no standard for that, so the only way is to implement the client-side code in javascript. It would be great if there was a standard for that so that all kinds of clients knew how to provide a proof of work, e.g. like this: WWW-Authenticate: Proof-Of-Work difficulty=5 challenge=XYZ Authorization: Proof-Of-Work abc Where sha256(abcXYZ) would have to start with at least 5 zeros.
- some_furry 1y agoWrite an RFC draft, toss it at the IETF. Seriously.
- dfawcus 1y agoThen have the server error response vend the Anubis JS as a fallback?
- GoblinSlayer 1y agohttps://github.com/vaxerski/checkpoint https://github.com/vaxerski/checkpoint
- prmoustache 1y agoI was thinking about adding a link to a page that is hidden in a one pixel image and same color as the page background. hiting it would mean a rule would be added on the firewall to ban that ip for a few weeks. The only is issue I can think of is there may be browsers or browser extensions that preload links to show thumbnails and users might be banned without knowing why.
- 8474_s 1y agoI've been seeing that anime girl pop-up in some websites, mainly because i use "rare" browsers.I prefer it over captchas and cloudflare "protecting websites from real traffic", whatever they're doing is just a few seconds and doesn't require solving captchas or something equally obnoxius like microsoft puzzles.
- kh_hk 1y agoClient side proof of work might be enough now but it won't last: solve challenge, reuse cookie. Ja4 fingerprinting is a new-ish in interesting approach, not for blocking but as an extra metric to validate trust on requests
- anonfordays 1y agoThis (Anubis) "RiiR" of haproxy-protection is easily bypassed: https://addons.mozilla.org/en-US/firefox/addon/anubis-bypass/ https://addons.mozilla.org/en-US/firefox/addon/anubis-bypass...