5 ms·
> ...we have tried to minimize the impact on real readers as much as possible. We have not gone with tools like Anubis, partly because it causes annoying delays
by harshreality 3mo ago
> ...we have tried to minimize the impact on real readers as much as possible. We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when you have millions of other people's machines to do the work on.
It's massively less annoying than a captcha, which is both a longer delay (typically, at present) and a massive cognitive distraction/roadblock.
The anubis author has stated they recognize it's an arms race, but PoW scales. Captchas and other signals are already at the end of the road; any additional difficulty increases false bot-positives, which are already unacceptably high.
For websites running dynamic languages, a binary (anubis is in go) sentry that operates before[1] the website is forced to expend any resources, is usually a large improvement over a site-hosted captcha. I would rather, and I think most humans would agree, have to wait a few seconds, maybe even closer to a minute in the future, to get a website access token good for a day or a week, than be forced to solve a captcha.
The dilemma for bots: when tokens are bound to the connecting ip, scrapers must limit the connecting IP pool for each site they want to scrape, becoming much more obvious and easy to block, or they have to use massive amounts of compute.
[1] this is true regardless of whether anubis is in reverse proxy mode or auth mode.
- Groxx 3mo agoAnubis is by far the least annoying throttler I encounter. Entirely agreed, just crank it up when you get a flood, I much prefer waiting a couple seconds to interacting with custom UI for tens of seconds. I'm so glad to see that (essentially) HashCash is coming back. Now we just need it for email, like it was originally designed for...
- deleted 3mo ago[deleted]
- alightsoul 3mo agoFrom my understanding this is also how cloudflare bot protection has worked for a long time, and then they look for entropy in user input to confirm the user is human. Also how recaptcha without images works.
- Rebelgecko 3mo agoCloudflare often just straight up blocks me or makes me do a captcha. IMO those are both much worse than Anubis
- alightsoul 3mo agoSounds like it's because your IP is on a data Center IP list
- Rebelgecko 3mo agoThat's lame of them, I'm on a residential plan. Maybe because I host my personal website from home?
- jazzyjackson 3mo agoGoogle and Cloudflare both are not just looking at entropy of mouse movements, that was cracked years ago, they are fingerprinting you and correlating your session with all your activity cross domains to score your botlike behavior.
- potamic 3mo agoI doubt they are doing it. You just have to get on a VPN to and see yourself being flooded with captchas despite browsing the web like a normal human and solving dozens of captchas along the way.
- alightsoul 3mo agoI guess that raises the bot score given that the vpn IP is a data Center IP and thus in a lot of ban lists
- Chu4eeno 3mo agoExcept when it throws you into a reload loop. It's pretty buggy, and trivial to bypass. And contrary to grandparent, PoW only worked because it was a novel thing to work around, a simple "type human" prompt would've worked as well. When anubis gets widespread enough users will still run the PoW in javascript or whatever while the scrapers will run much more optimized native code, so no, it doesn't scale.
- harshreality 3mo agoPutting aside the question of whether it will continue to work, even somewhat, against botnets, I find your first paragraph confusing. Reload loops, or being able to "bypass" anubis (unless you merely mean bypassing it for the token validity period by solving a challenge), sound like misconfigurations. There's no reason for anubis itself to cause reload loops; it's tricky to configure a webserver to use it in some scenarios. Any ability to bypass anubis probably means the site is using it in auth/challenge mode only, and then misconfigured their webserver's auth checking. Or it's a bug. If you mean the double-spend tavis mentioned in his blog post which previously made the HN frontpage, that was patched right after it was reported to the maintainer almost a year ago.
- deleted 3mo ago[deleted]
- Gigachad 3mo agoI looked this up and realised it’s the page I’d seen briefly on a range of websites lately. It’s not annoyed me at all. Not nearly as much as having to complete captchas with slow refreshing tiles.
- inigyou 3mo agoA fun fact about Google captchas: they've often decided whether you will succeed or fail the captcha before you do the captcha.
- RickHull 3mo agoThis seems nonsensical. Care to elaborate?
- inigyou 3mo agoIf Google determines you're an undesirable user, doing the captcha is just an exercise to waste your time.
- boramalper 3mo agoI certainly experienced this (the vicious try-again cycle) but curious if you have any sources for this?
- applfanboysbgon 3mo agoA person's testimony is a source. I can add mine: you can tell when you're truly blocked because if you click for the accessibility audio-based captcha it will actually tell you you're blocked (but, if you did the visual captcha, would simply loop forever while telling you you did it wrong). I don't think you'll find an article by Google saying "yes, we sometimes completely block users while making it look like they're not blocked and wasting their time".
- herpdyderp 3mo ago
- miyuru 3mo agoFor me Cloudflare is worse, it takes more than 5 seconds, where as anubius take 1-2 secs. funny with all the IP information they have, cloudflare cannot do a better job. (I am on IPv6) and most of the time, its on marketing product pages like in framework main site, which can be cached.
- freehorse 3mo agoImo the worst is recaptcha. At least with cloudflare the work you have to provide is minimal. With recaptcha it can take me much longer than 5 seconds, and lately I have trouble even completing their challenges correctly. Nowadays if I see a (recaptcha) captcha I drop the site unless I must visit it for some reason, it is not worth the time, the effort or the annoyance.
- miki123211 3mo agoMost CF / Recaptcha problems are users going "off the golden path", and not realizing that their config changes are at fault. If you're on a consumer router, using a mainstream stock browser with stock settings (maybe plus uBlock Origin), with your Google account logged in, it's very, very likely to just work. If you're part of the .01% of users with opinions about that sort of thing... you're not worth optimizing for.
- aiiotnoodle 3mo agoA lot of users also run cheap cracked fire sticks and other low reputation hardware that's proxying their residential traffic for nefarious reasons which makes all the big providers put up their guard.
- freehorse 3mo agoAt least for me, CF is fine; recaptcha is the only one I really have problems with. I dont care what recaptcha wants to optimize for. I dont think that using a vpn is that a rare thing anyway. If others have figured out how to do it without requiring spending 30 seconds to solve a captcha, I dont see why websites still use recaptcha/captchas for that. And that it is "my fault" not being logged into google I was least expecting to see here.
- Aurornis 3mo ago> just crank it up when you get a flood, A few months ago there was a story posted here about someone who completely eliminated crawlers on their website with Anubis. I think it was getting upvoted before users were clicking the article because if you did, you had to leave the Anubis PoW page open for several minutes before you could get into the site. The Anubis difficulty scale is unintuitive and the difference between a small delay and becoming unusable is easy to cross.
- m463 3mo agoAt least anubis works for me. (I run umatrix) Unfortunately whatever HN is using routinely blocks my login with "Sorry." some websites just always give me 403.
- duskwuff 3mo ago> Unfortunately whatever HN is using routinely blocks my login with "Sorry." I believe that's the HN application itself, not a WAF in front of it.
- asdfsa32 3mo agoHN is surprisingly very very guilty of a whole lot of anti-user patterns and behaviour that other companies get regularly lamented. Poor accessibility, bad mobile support, no options to delete content beyond a narrow window.
- hurfdurf 3mo agono options to delete content beyond a narrow window. Good.
- asdfsa32 3mo agohurfdurf says it is okay to keep material online forever. It is easy to say that when you publish as hurfdurf.
- ambigious7777 3mo agoi personally like using http://hcker.news http://hcker.news as a reader, its much nicer
- kps 3mo agobad mobile support Good.
- nkrisc 3mo ago
- PlasmaPower 3mo agoI don't think PoW scales, because if the bot authors get serious they'll start using native implementations that are much more efficient than the web ones real users are running. In theory maybe Anubis could start using WebGPU to help close that gap, but then anyone without WebGPU support is out of luck. Then again, a large portion of the problem seems to be bots making way too many requests and in general not being optimized in the first place, and this does help filter those out.
- ammario 3mo agoThere are PoW approaches that even the playing field between data centers and desktops. RandomX is my favorite.
- mootothemax 3mo agoInteresting. How do they tell the difference between legitimate and forged ip owner records?
- anonym29 3mo agoIt's not about traffic identification at all, but rather a hashing algorithm that is deliberately resistant to parallelization and GPU/ASIC acceleration, which shrinks the gap in solving speed between the fastest systems (i.e. datacenter-class compute resources) and typical systems (e.g. the CPU in your smartphone or laptop).
- Sesse__ 3mo agoUh, is it resistant to parallelization across multiple sites? Because that's the situation for the scrapers. They're not trying to solve a single PoW challenge across many cores.
- anonym29 3mo agoTypically, the machine doing the content processing, including solving PoW, is the centralized "control" node described in the article, not the machines who's IP addresses are being used. In typical residential proxy networks, the residential proxies are exposed to the customer (the person paying for and using the proxies) as just SOCKS5 addresses, and no computational power from those compromised devices is made available for the scraper besides that used to power the SOCKS5 server itself, the customer is just paying for the transport and address (and indeed, is often billed on either a per-GB or per-IP basis). In effect, if the customer (the entity paying for and using the proxies) wants to solve PoW challenges through those connections, it is indeed the customer who must pay that compute cost, not the compromised devices. Note that this is the case for a majority of, but not all, residential proxy networks, which often are built through quasi-voluntary distribution channels, including SDKs included in otherwise legitimate mobile applications distributed through Apple's App Store and Google Play. These distribution channels tend to be categorically unavailable (or at least unreliable) for true RAT-style malware that enables remote operators to dynamically assign arbitrary computational workloads to client devices. This isn't to say that true botnets built with actual malware delivered through either software exploits, phishing attacks, or watering hole attacks don't also perform as residential proxy networks, but such categories are a relatively small subset of all residential proxy networks, and there are much higher ROI malicious activities to be performed on these devices rather than serving as relatively mundane traffic networks for scraping.
- noncoml 3mo agoOr you can go full Reddit and just block anything that seems even remotely suspicious. Your sibling, roommate, neighbor that uses your internet, previous IP owner, posts too much? You get blocked too. Using VPN? Blocked. Your iPhone is too old, blocked. Your screen brightness too low? Believe or not, blocked.
- MikeRichardson 3mo agoParks and Rec reference?
- lsaferite 3mo ago> Your screen brightness too low? Believe or not, blocked. ... What?!
- rwmj 3mo agoEven worse, not blocked, shadowbanned.
- ValdikSS 3mo agoMy forum got scraped so hard that the ISP blackholed the IPv4 several times a week. I've ended up putting only IPv6 on the domain. It's running this way for 2 years already.
- pineapplepizza6 3mo agoIt's fun to log in with a banned Reddit account on a shared IP, which bans every other Reddit login on the same IP.
- leptons 3mo agoI quit reddit because of all of this nonsense. After 15 years on reddit, my life has been much better for quitting it. Reddit is a cesspool.
- deleted 3mo ago[deleted]
- mootothemax 3mo ago> The anubis author has stated they recognize it's an arms race, but PoW scales. The scraper wars are largely between script kiddies and people with both deep intimate networking and DOM knowledge. Yes greyhairs, I’m looking at you. The problem is, you can’t PoW every page load and resource request because the user experience will suck and people will run away. And that window - the gap between what people will tolerate vs draconian enforcement - is exactly what the scrapers exploit. And looking at the PoW options out there - I’ve seen at least one PoW WAF (honestly can’t remember if azure or amazon) have their PoW boil down to repeated trigonometric functions, ie very optimisable. It’s a neat concept, but the answer and future to my eyes look bleak.
- harshreality 3mo agoAnubis's default 1-week token lifetime may not be nearly enough to dissuade enough scraper networks to make a difference, particularly with the default weight->difficulty level hierarchy, but that's for individual site admins to determine. We can all argue based on how we envision "ideal" scraper networks being run and whether the web-PoW concept would stand up to that. However, what matters at present is that anubis helps many sites cope with misbehaving bot scrapers written by the script kiddies you mention, who don't care if the internet burns as long as they finish their scrape 1 hour faster. If anubis motivates them to devote a few brain cells to make their scrapers smarter, they may also fix the scrapers to not take down the sites they're scraping.
- inigyou 3mo agoOne of the ideas behind Anubis was to incentivize a scraper to stop hiding, because every change of identity brings another challenge page.
- miki123211 3mo agoOh, but you can PoW every page. Your typical end user doesn't switch IPs that often, so it's fine to Anubis them again when they do. A scraper, on the other hand, has a tradeoff to make between rotating ips often (requiring a challenge on every request) or keeping only a few IPs (making cross-request identification much more valuable and reliable).
- jimmydorry 3mo agoPoW barely affects the "residential proxies" aka. malware botfarms. The IPs are free for them and siphoning additional system resources for PoW doesn't matter at all for them. PoW only affects the large centralised scraping by the AI providers, which are not operating behind "residential proxies".
- anonym29 3mo agoMost users of residential proxies just get a SOCKS5 address and routing, they don't actually get computational resources of the infected systems beyond that. The user of the proxies, the operator of what the article describes as a control node, would be the device responsible for the PoW. Do you have any evidence that AI providers aren't using residential proxies?
- jimmydorry 3mo agoYes, the status quo right now is merely bandwidth... but if there was money to be made in providing a small amount of compute for the sake of solving the next gen captcha's, you can bet offerings will expand to meet it. It's impossible to prove a negative. They could all be running secondary scrapers using malware proxies... but what we do have is plenty of evidence of them using fixed IP pools with appropriate user agents. I don't see any Chinese user agents... so two guesses who may be driving the bulk of these AI scraper requests via residential proxies.
- inigyou 3mo agoResidential proxy bandwidth is extremely expensive, comparatively speaking. It can be up to $1 per GB but is more typically about $0.20 per GB.
- jimmydorry 3mo agoRight, this is orthagonal to the discussion though. While the IPs and bandwidth might be free, managing the malware botnet and trying to keep a low profile so as to not attract attention of authorities, makes it a risky market to cater to. Those rates are still cheaper than some datacentres charge in parts of the world.
- armchairhacker 3mo agoPoW can theoretically scale effectively infinite because it can mine cryptocurrency. Millions of compromised IoT devices hitting your server? Now you have enough money for a faster server. It doesn’t matter that the challenge must be verified: present multiple challenges, some are verified while others mine crypto.
- Sesse__ 3mo agoThis is called “installing a cryptominer on your web page” and is generally considered illegitimate.
- armchairhacker 3mo ago> generally considered illegitimate But why? Obviously an unjustified cryptominer is bad, like unnecessarily slow JavaScript, but this one has a good purpose and to the user is no different than PoW.
- geocar 3mo ago> but PoW scales Not if the honest party is doing it in a browser: The same computer can so any POW so much faster in C than any amount jf JS and WASM that it will never ever ever be a contest. > becoming much more obvious and easy to block, or they have to use massive amounts of compute. If you believe this, please contact me: I think compute is free[1] and can probably help you out. [1]: https://news.ycombinator.com/item?id=30175269 https://news.ycombinator.com/item?id=30175269
- swinglock 3mo agoCan you not design a PoW that is most efficient in a browser? Don't brute force hashes like Hashcash/Bitcoin, do something similar to RandomX instead but in JS. Browsers ought to run the fastest JS interpreters already so if interpreting JS becomes the bulk of the work, that attack might not work. Maybe even involve the DOM or whatever else makes sense.
- deleted 3mo ago[deleted]
- geocar 3mo ago> Browsers ought to run the fastest JS interpreters already Well they don't. Users want the website to work sooner, and care little about whether a for-loop of elements take 10ms or 20ms if it only happens once. JS can be AOT compiled if you can wait a few _seconds_ -- which users don't want, so browsers don't bother. Our attacker however, rightly observes they only have to pay that compilation cost once.
- swinglock 3mo agoIf it takes seconds to compile it then the browser wins. The PoW only runs for seconds already.
- geocar 3mo agoNo no no. You don’t realise the compiled intermediate can be _trivially_ reused for every request, and identified by the hash of the JavaScript code. Browsers already do this, it’s just your attacker has a bigger cache, and not by a little bit. Browsers _intentionally_ limit the amount of memory and cpu cycles a tab can consume to make for a nicer human experience.
- iiiwio 3mo ago[flagged]
- jsnell 3mo agoProof of work does not scale. It trades something fungible and incredibly cheap (CPU) for something incredibly expensive (user-visible latency). There is no set of parameters where the cost is going to be a meaningful deterrent to any kind of abuse (even something as low-yield as scraping) without adding crippling amounts of latency to real users. > The dilemma for bots: when tokens are bound to the connecting ip, scrapers must limit the connecting IP pool for each site they want to scrape, becoming much more obvious and easy to block, or they have to use massive amounts of compute. There is no dilemma. They get a token, they maybe do some automated multi-armed bandit per-site to figure out how to maximize the extraction rate they get from a single token, and then they use an IP for that many requests / that amount of time before ditching it.
- out_of_protocol 3mo ago> It trades something fungible and incredibly cheap (CPU) it could be RAM-bound, which is very much NOT cheap nowadays :)
- dannyfritz07 3mo agoYes, but the people with the RAM nowadays are the data centers, not the end users.
- deleted 3mo ago[deleted]
- crote 3mo agoSure, but end users as a group still have a significant amount of RAM. Even on a low-specced machine you can afford to have the currently-active website tab consume a few hundreds of megabytes of RAM. It was mostly sitting idle anyways, so most people aren't even going to notice. Multiply that by a few thousand concurrent visitors and you're burning hundreds of gigabytes without anyone caring. Some scraper being forced to install hundreds of gigabytes of extra RAM in their crawling node? They will notice having to go from a $0.50 / h instance to a $15.00 / h one, solely for the extra RAM requirement.
- mschuster91 3mo ago> The dilemma for bots: when tokens are bound to the connecting ip, scrapers must limit the connecting IP pool for each site they want to scrape, becoming much more obvious and easy to block, or they have to use massive amounts of compute. You can't do that any more. Too many ISPs, especially mobile carriers, don't hand out anything resembling a fixed IP address any more. It's CGNAT and constantly changing IP addresses alllll the time now.
- pocksuppet 3mo agoOn a small enough site (even LWN might qualify) the chance of two random sets of client IPs intersecting can be quite low. Private trackers do this. If they ban a user that geolocates to a certain city and ISP, they'll ban new signups from that city and ISP because there's probably only a few users from the same city and ISP. And then report to their friends at other trackers, that a user with that city and ISP is trying to evade a ban.
- setupminimal 3mo agoWell, we don't use a captcha either. If it were a choice between a captcha and a proof of work system, we'd have to reevaluate things. Luckily, for now, we're able to get away with a much lighter touch.
- aftbit 3mo agoI'm not gonna wait a minute to read an article. Instead, I'll either just leave, or go query it from archive.$tld that bypasses it for me.
- dlenski 3mo agoAnubis appears to be a temporarily-useful stopgap that has been cargo culted into prominence and an expectation of permanent usefulness, for reasons I don't fully understand. The cost of solving the default Anubis PoW is negligible on cloud servers, and it's even lower if you use native code rather than JavaScript to solve it, which Tavis Ormandy helpfully demonstrated last year (https://lock.cmpxchg8b.com/anubis.html https://lock.cmpxchg8b.com/anubis.html). If Anubis were to be even more widely adopted, botnet operators would surely adopt and optimize native code solvers en masse. So Anubis doesn't do much to stop bots, but it makes otherwise lightweight websites (little JavaScript or interactivity) almost unusable on low-resource systems like my old phone or an old Atom-based nettop. > when tokens are bound to the connecting ip, scrapers must limit the connecting IP pool for each site they want to scrape This "IP-bound proof-of-work" thing is gonna kill multipath TCP and bring down IPv6 with it. Uffff.
- amlib 3mo ago> If Anubis were to be even more widely adopted, botnet operators would surely adopt and optimize native code solvers en masse. Then anubis adopts it itself, increases the amount of work that needs to be done and the bar stays the same again for everyone? Seems like mostly a non-issue unless there is an arms race towards ever more optimized solvers which I don't believe is possible.
- dlenski 3mo ago> > If Anubis were to be even more widely adopted, botnet operators would surely adopt and optimize native code solvers en masse. > Then anubis adopts it itself, increases the amount of work that needs to be done and the bar stays the same again for everyone? No, the bar doesn't "stay the same" for everyone interacting with Anubis. My otherwise-perfectly-usable 8-year-old phone, which can't be patched to run a native solver, becomes even more unusable on sites gates with proof-of-work challenges like Anubis. This is the whole problem with PoW. It forces thousands or millions or billions of client devices to do increasing amounts of useless work which is relatively easy for cloud-based attackers to adapt to, but very difficult for hardware-constrained and software-ossified mobile clients to adapt to. In other words, it asymmetrically punishes the clients that it's not intending to punish.