5 ms·
Looks cool. But please help me understand. What's to stop AI companies from solving the challenge, completing the proof of work and scrape websites anyway?
by throwaway150 1y ago
Looks cool. But please help me understand. What's to stop AI companies from solving the challenge, completing the proof of work and scrape websites anyway?
- perching_aix 1y agoNothing. The idea instead that at scale the expenses of solving the challenges becomes too great.
- crq-yml 1y agoIt's a strategy to redefine the doctrine of information warfare on the public Internet from maneuver(leveraged and coordinated usage of resources to create relatively greater effects) towards attrition(resources are poured in indiscriminately until one side capitulates). Individual humans don't care about a proof-of-work challenge if the information is valuable to them - many web sites already load slowly through a combination of poor coding and spyware ad-tech. But companies care, because that changes their ability to scrape from a modest cost of doing business into a money pit. In the earlier periods of the web, scraping wasn't necessarily adversarial because search engines and aggregators were serving some public good. In the AI era it's become belligerent - a form of raiding and repackaging credit. Proof of work as a deterrent was proposed to fight spam decades ago(Hashcash) but it's only now that it's really needed to become weaponized.
- marginalia_nu 1y agoThe problem with scrapers in general is the asymmetry of compute resources involved in generating versus requesting a website. You can likely make millions of HTTP requests with the compute required in generating the average response. If you make it more expensive to request a documents at scale, you make this type of crawling prohibitively expensive. On a small scale it really doesn't matter, but if you're casting an extremely wide net and re-fetching the same documents hundreds of times, yeah it really does matter. Even if you have a big VC budget.
- charcircuit 1y agoIf you make it prohibitively expensive almost no regular user will want to wait for it.
- bobmcnamara 1y agoExponential backoff!
- xboxnolifes 1y agoRegular users usually aren't page hopping 10 pages per second. A regular user is usually 100 times less than that.
- pabs3 1y agoI tend to get blocked by HN when opening lots of comment pages in tabs with Ctrl+click.
- xboxnolifes 1y agoYes, HN has a fairly strict slow down policy for commenting. But, that's irrelevant to the context.
- pabs3 1y agoI meant to say article pages not comment pages, but ack.
- Nathanba 1y agoYes but the scraper only has to solve it once and it gets cached too right? Surely it gets cached, otherwise it would be too annoying for humans on phones too? I guess it depends on whether scrapers are just simple curl clients or full headless browsers but I seriously doubt that Google tier LLM scrapers rely on site content loading statically without js.
- ndiddy 1y agoThis makes it much more expensive for them to scrape because they have to run full web browsers instead of limited headless browsers without full Javascript support like they currently do. There's empirical proof that this works. When GNOME deployed it on their Gitlab, they found that around 97% of the traffic in a given 2.5 hour period was blocked by Anubis. https://social.treehouse.systems/@barthalion/114190930216801561 https://social.treehouse.systems/@barthalion/114190930216801...
- dragonwriter 1y ago> This makes it much more expensive for them to scrape because they have to run full web browsers instead of limited headless browsers without full Javascript support like they currently do. There's empirical proof that this works. It works in the short term, but the more people that use it, the more likely that scrapers start running full browsers.
- deleted 1y ago[deleted]
- sadeshmukh 1y agoWhich are more expensive - you can't run as many especially with Anubis
- SuperNinKenDo 1y agoThat's the point. An individual user doesn't lose sleep over using a full browser, that's exactly how they use the web anyway, but for an LLM scraper or similar, this greatly increases costs on their end and partially thereby rebalances the power/cost imbalance, and at the very least, encourages innovations to make the scrapers externalise costs less by not rescraping things over and over again just because you're too lazy, and the weight of doing so is born by somebody else, not you. It's an incentive correction for the commons.
- ronsor 1y agoI know companies that already solve it.
- creata 1y agoWhy is spending all that CPU time to scrape the handful of sites that use Anubis worth it to them?
- vhcr 1y agoBecause it's not a lot of CPU, you only have to solve it once per website, and the default policy difficulty of 16 for bots is worthless because you can just change your user agent so you get a difficulty of 4.
- wredcoll 1y agoI mean... knowing how to solve it isn't the trick, it's doing it a million times a minute for your firehose scraper.
- udev4096 1y agoAnubis adds a cookie name `within.website-x-cmd-anubis-auth` which can be used by scrapers for not solving it more than once. Just have a fleet of servers whose sole purpose is to extract the cookie after solving the challenges and make sure all of them stay valid. It's not a big deal
- fc417fc802 1y agoRequests are associated with the cookie meaning you can trace and block or rate limit as necessary. The cost of solving the PoW is the cost of establishing a new session. If you get blocked you have to solve again.
- userbinator 1y agoThis is basically the DRM wars again. Those who have vested interests in mass crawling will have the resources to blast through anything, while the legit users get subjected to more and more draconian measures.
- SuperNinKenDo 1y agoI'll take this over a Captcha any day.
- userbinator 1y agoCAPTCHAs don't need JS, nor does asking a question that an LLM can't answer but a human can. Proof-of-work selects for those with the computing power and resources to do it. Bitcoin and all the other cryptocurrencies show what happens when you place value on that.
- fc417fc802 1y agoYou can provide visitors a choice. > Your visit has been flagged. Please select: Login, PoW, Cloudflare, Google.