16 ms·
Creepy Crawlies
- jruohonen 1mo agoOff-topic, but anyone with which he did the plots?
- dingaling911 1mo agolooks like this https://github.com/rfonseca/xkcd-gnuplot https://github.com/rfonseca/xkcd-gnuplot
- deleted 1mo ago[deleted]
- jruohonen 1mo agoThanks, and, yes, I don't write with LLMs, as seen above.
- daveguy 1mo agoI would rather see some awkward phrasing than bland LLM slop! Also, many plotting libraries include an xkcd style these days: https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.xkcd.html https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.... And libraries for various languages: https://github.com/timqian/chart.xkcd https://github.com/timqian/chart.xkcd So if you have a preferred dev environment there's probably a way to set it to xkcd style.
- electrogas 1mo agoor this: https://matplotlib.org/stable/gallery/showcase/xkcd.html https://matplotlib.org/stable/gallery/showcase/xkcd.html
- Artoooooor 1mo agoHow expensive would AI access be if every user paid their fair share instead of shoving it on the people doing the actual work?
- parineum 1mo agoYou mean shoving it onto the investors?
- inigyou 1mo agoNo they mean the people doing the actual work, I think.
- parineum 1mo agoThe people doing the work aren't paying for AI, they're getting paid by AI. The people who are having the costs of users not paying "their fair share" (this phrase has officially jumped the shark) are the investors who are subsidizing these companies.
- wredcoll 1mo agoSee also: the price for uber rides.
- initramfs 1mo agoI've been noticing page views in the past several months with a much wider span of origin on my Blogger stats. Before I would get a few from several countries, but now I am getting views from tiny countries and obscure or outdated browsers and operating systems, which leads me to think scrapers could be using VPN services in various countries along with header anonymizers that mask the device that they are using. Extensions like ModHeader, BrowserMask do this: https://github.com/apify/crawlee-python https://github.com/apify/crawlee-python https://github.com/mthcht/Masquerade-Spoofer https://github.com/mthcht/Masquerade-Spoofer Great for AI scrapers, bad for hosters and everyone else.
- chuckadams 1mo agoGiven the nature of git, wouldn't all that HTML be highly cacheable? I get that's not free either, but it's got to be a lot less intensive than having cgit generate it every single time.
- Zariel 1mo agoThat was my first thought, varnish (vinyl these days) in front of the website should dramatically reduce this as the git repo should be practically static for most of the content.
- teo_zero 1mo agoBut there are "cubic bazillions" of possible URLs that are being requested. Even if they boil down to "only" some millions actual commits, their rendered HTMLs are all different.
- kijin 1mo agoYes, this is very difficult to solve for sites with many URL variations, like git repos and heavily threaded forums. Your cache is always full, but the hit ratio is abysmal.
- oowa 1mo agono its only difficult for you. and only at this moment... any minute now you will see the way. btw, the OP is just saying "its OK for now". And the OP is just telling us: this is what has been happening... maybe its difficult for OP also, but they didnt say that. they just said their current setup cant handle it. old tech...have a hackathon to solve for this. OP says it's not a problem for him right now but if we extrapolate what he's talking about it's definitely a problem aaaaaaand It's totally solvable, Even with all of the crazy combinations he's talking about it's still solvable. And it's already been solved using patterns we see in streaming services. This is completely hackathonable. but why do we even need to bother with this? The slurpers are the cause of this, and they can cause this problem because of Murphy's Law. well you can only account for Murphy's Law with good architecture or something like that or whatever. Ha ha hackathon.
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- lkbm 1mo ago> So, you'd think that something that pretends to be “Artificial Intelligence” would use the most efficient way of using our data for training purposes, right? Clone the repos, walk every commit. Done. If repos like this were 10%+ of what they crawled, having a special case for clone-able repos would be smart, but if you're crawling everything, you're not going to do an efficiency tweak for each special case that has a more parse-able option.
- phmx 1mo agoI guess GitHub is in a similar bunch of sources, it should be also more efficient to crawl by cloning. Anyway, isn’t it the whole sales pitch that it generates tailored solutions fast?
- kalkin 1mo agoThis tick about "if AI smart how come crawler dumb" is in most complaints I've read about AI crawlers and I've started to find it pretty annoying. The crawlers might be written using AI but they're evidently not actually running AI inference over the pages they get back--besides being able to tell this from the behavior, if this is pretraining input, that's enormous scale, so it'd mean a large increase in effective training cost. Naively assume inference costs are equal to pretraining costs (probably not true but maybe right order-of-magnitude) and it's a doubling. This ends up tacitly turning a very legitimate complaint (ill-behaved crawlers) into a justification for head-in-the-sand AI denialism.
- bigstrat2003 1mo agoThere's nothing "denialist" about recognizing the utter stupidity of systems that are being mislabeled as "AI". You judge a tool by its results, and the results have been very poor indeed. The only heads in the sand are those whose owners continually refuse to recognize the proofs before their very eyes that there's zero intelligence here.
- kalkin 1mo ago> zero intelligence here That's head-in-the-sand stuff. AI is certainly very capable of being dumb (as are humans). But: > “The problem was in need of a new real idea, which this new result seems to provide,” says James Maynard, a mathematician at the University of Oxford. “It seems that the AI has made a genuinely interesting mathematical contribution.” https://www.scientificamerican.com/article/no-ai-didnt-just-solve-the-thorniest-problem-in-math/ https://www.scientificamerican.com/article/no-ai-didnt-just-... Nobody a decade ago would have said "oh yeah solving a bunch of open problems in research mathematics, and finding a bunch of zero days in Chrome and Firefox, and winning literature prizes, are things that don't require intelligence."
- nicman23 1mo agocouldn't you have anubis on a dynamic difficulty? ie if a ip requests more than 1k pages per day +1 the difficulty ?
- dunder_cat 1mo agoYes, but the article (not to call you out - I just think it's a very important point!) points out that this type of throttling would not be effective: > Suddenly, the crawlers were coming from millions of random residential or mobile IPs, all pretending to be random modern browsers. An IP like that would make 4-5 requests and then never show up in the logs again. There was no point in banning them, because by the time you figured out that they were bots, they were already done with you. You just needlessly ballooned your firewall ruleset by adding IPs that would never be back. Without something like cookies (which are almost certainly tossed after the IP is rotated) or some other persistent identifier, you are stuck have to apply mitigations that scale with the load you're encountering, which means longer challenges for everyone or degraded functionality, like removing some of the fancier cgit features.
- marginalia_nu 1mo agoI've had a fair bit of success with increasing the bot mitigation based on a global rate limit. During periods of high request rates, I throw progressively more hurdles at the bots, and during periods of low request rates I disable them all.
- nicman23 1mo agoyeah i my head i thought they meant 4-5 _K_ requests
- PinkaDunka 1mo agoMaybe anubis difficulty should depend on the age of commit. This year - 4, everything older 8
- nicman23 1mo agoor on cache hit /miss
- Velocifyer 1mo agoBut why don't they just git clone?
- lkbm 1mo agoBecause they're crawling a billion webpages, only a tiny fraction of which can be git cloned, and configuring a special case just for that tiny fraction isn't worth the effort (of the crawlers).
- acedTrex 1mo agoBecause the crawlers dont care, they are the internets parasites. Their creators care nothing for people or systems downstream of their greed.
- rcxdude 1mo agoThese crawlers (in contrast to e.g. googlebot and similar better behaved crawlers) are not very smart: they seem to make very little effort to avoiding crawling useless deep trees of generated pages. About the only thing they seem to put a lot of effort into is avoiding blocking. (I'll note that while these are generally attributed to AI data gathering because of the timing of when they took off, it's not actually obvious who's running these bots. The big players all have crawlers that identify themselves and are reasonably well behaved, but I don't know if anyone has managed to positively attribute these other ones to any particular group)
- voakbasda 1mo agoWhy do we think that only “good guys” are training LLMs? I imagine organized crime is getting in on the game too.
- ipdashc 1mo ago> it's not actually obvious who's running these bots This is fascinating to me. It's a large enough phenomenon that it's affecting the entire Internet and yet nobody seems to know yet who's actually doing it. Which isn't surprising, of course, it's hard to trace back to a source through all these proxies and it's probably a bunch of distinct groups anyways, but still! Personally I have to wonder how much of it is "scrapers for training data" vs just tool-use LLMs. Even if you use chatgpt in thinking mode you can clearly see it searching and visiting a bunch of different websites to answer a question, presumably faster than any human would. That's got to add up. It's got me wondering why everyone seemingly discounts that as an option
- AshamedCaptain 1mo agoThis is bad enough that I'm going to stop serving cgit. I've been doing cvsweb, then subversion, then cgit over my home server for many many years and for the first time ever this is annoying my own bw usage. It's ridiculous also how you ban an IP then 1 second later another one picks up from where the first one left on.
- rwmj 1mo agoI had to take down my cgit repository a few months ago. The load was causing other VMs on the same machine to become unusable.
- singpolyma3 1mo agoI've been using git-arr instead with some success
- a-dub 1mo agoi wonder what they're all up to. i imagine some are scraping datasets for pre-training, others are probably real-time scrapers looking for security bugs, even more still are agents working on coding tasks and looking at the kernel. also interesting to think about solutions: does everything need to be optimized now for weird access patterns that proliferated ai creates? do the ais need to have behavior trained in to be better netizens? is this the end of anonymous browsing and the beginning of an era where one has to attach an identity to all requests? or the end of community hosted free information services more broadly?
- yellow_lead 1mo agoHigh Anubis difficulty is annoying the hell out of me for several sites. And it's starting to not block LLM bots anymore? > 33% are now solving the math and getting through to the main site — because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge.
- sethops1 1mo agoIn a few years the VC money will dry up and this gross overspend on slurping data will end.
- igor47 1mo agoVisions of vast data centers surrounded by fields of browning grass, in which aging, rusting, formerly extremely expensive hardware is spending billions of compute cycles looking at anime catgirls
- rpcope1 1mo agoSounds like an even shittier version of the Lorax. :(
- pixl97 1mo agoUnfortunately we're apt to run into some kind of Jeavons Paradox where the hardware gets so much faster in those few years will be able to slurp massive amounts of data cheaply so the problem never really ends.
- marginalia_nu 1mo agoIt's very likely the last few years of bot behavior is the consequence of the residential proxy business booming. This is indirectly due to AI company crawling, but the fact that they are as cheap and available as they are changes the incentives for anyone using them toward reckless and unsustainable request behavior, as there is no risk of burning your IPs, and very small chances of seeing any consequences of essentially DDoS:ing a website.
- Demiurge 1mo agoI maintain a formerly popular gaming website, and it used to have hundreds of legitimate requests per second. The load would be especially high during popular event times. So, it’s always been running on a dedicated server. It also has an “online users” counter, which attempted to count real user sessions of unauthenticated user which still maintained a session, which lets them comment, or modify certain filter and display options. It never counted the Google bot. Over the last few years this counter went from 100-200 users online to thousands. I have been very hands-off with it for many years, doing minor upgrades and backups. However, the site also has gotten quite slow, these sessions were obviously impacting it. So, I finally investigated these crawlers, and yes, it turns out it’s an insane amount of traffic that is entirely artificial, the site has just a handful of real users, and thousands of these crawling sessions that actually try to do everything they can, click every button. It doesn’t help that sort and search were implemented using GET links. I fixed the counter to exclude the crawlers, but I have a bit of a dilemma. I don’t want to stop the bots from updating their knowledge based on all the content. The best solution I could find is the new CloudFlare feature where they might charge the crawlers for every request, or otherwise block them. I think that’s a fantastic idea for the internet, at large. I signed up for the beta access, but haven’t heard from them again. I do think it’s unfortunate that this requires CloudFlare and the middleman. Overall, it seems like the LLM are really straining the internet economy, the openness of it. Email spam used to be the worst, but the organized trillionaire labs sucking up the entire internet is going to break something if we don’t preempt them better. It’s too bad the copyright and public internet systems are not acting quick enough. And I think there is no reason to act like this race really has to be at such a breakneck speed.
- andai 1mo ago> And I think there is no reason to act like this race really has to be at such a breakneck speed. I would promote this idea to all of my competitors. Nah mate, you don't have to ask her out right now. You can wait until next week ;) Anyway, silver linings, looks like we're finally going to get widely adopted infra for microtransactions. https://web.archive.org/web/20030202042510/http://www.openp2p.com/pub/a/p2p/2000/12/19/micropayments.html https://web.archive.org/web/20030202042510/http://www.openp2...
- acedTrex 1mo agoIt feels inevitable that many systems will have to go to a login/trusted ip source type system. Its just not feasible to continue to operate with 99% of your traffic being fake.
- GoblinSlayer 1mo agoThey go to cloudflare.
- petesergeant 1mo ago> Training an LLM on content produced by the LLM gives it the equivalent of a digital prion disease Is it foolish of me to have expected more from a blog post on kernel.org?
- theandrewbailey 1mo agoAre you trying to say that's bad writing? I think it's a good metaphor for a documented phenomenon: https://en.wikipedia.org/wiki/Model_collapse https://en.wikipedia.org/wiki/Model_collapse
- Lerc 1mo agoAs the article states, this phenomenon may be documented, but there is no consensus that it describes any practical reality. The predicted consequences have now had time to manifest, and have not done so. This makes the claim either false or overstated. Perhaps there will be issues in the future, but to date there have been many claims that AI development will stall (for a variety of reasons). If they were the critical weaknesses they have been portrayed as, models would not have advanced to the level they are today. If you have a hypothesis, make a clear prediction based upon it. If you start pushing the date forward after each failed prediction, you end up looking like a hapless doomsday cult. If your hypothesis is correct however, your prediction should actually happen. Then provided you have not made so many predictions to get one right by chance, people will take what you have to say seriously.
- desterothx 1mo agoYou state it as if its only the quality of the hypothesis that matters, but you are ignoring an important part of it, timing. During the 08 financial crisis Burry had a hypothesis that was correct, however he almost went bankrupt still because he thought it would happen earlier than it did because of the government bailouts. He was pushing the day forward, and was looking like a "hapless doomsday cult". His hypothesis still turned out correct
- 1mo ago
- easton 1mo agoSide note: why are shallow clones evil? I always thought they were cheaper, but I guess that’s really just for my disk space. (since the server has to compute what blobs to give you instead of just “everything”?)
- jacobvosmaer 1mo agoNormal clones can reuse delta-compressed data the server stored on disk. Shallow clones impose a negative constraint: do not transfer data outside the requested commit depth. Pre-computed delta chains than contain unrequested data become unusable and the server must do delta compression on the fly to satsify the shallow clone.
- Hackbraten 1mo agoWhat I find remarkable is that for at least a decade, i.e., long before LLM scrapers were a thing, GitHub engineers have been reaching out to popular package manager projects, asking them to do away with shallow clones [0] [1]. They basically used the same reasoning as your comment did. [0]: https://github.com/Homebrew/brew/pull/9383 https://github.com/Homebrew/brew/pull/9383 [1]: https://github.com/CocoaPods/CocoaPods/issues/4989#issuecomment-193772935 https://github.com/CocoaPods/CocoaPods/issues/4989#issuecomm...
- crote 1mo agoSo why not introduce a semi-shallow clone option? If I'm doing a shallow clone it isn't because I only want to receive a specific commit, it's because I don't want to burn a giant amount of disk space and network traffic on a full history. In most use cases it would be perfectly acceptable for the server to send additional data. The client doesn't care about it because it is meaningless to them, but if it results in a significant load reduction on the server's side they don't really mind receiving it either. A 100MB shallow checkout coming with 400MB of garbage still beats cloning an entire 5GB history!
- ygouzerh 1mo agoIt's the point that surprised me the most! We always used shallow clones, to speed the CI, I didn't knew that it got that much impact server side!
- feelamee 1mo agoHm, interesting - how will it look the actual solution for such problems in the future. I suppose the issue will continue to grow. First idea - there should be some cost for sending traffic somewhere. And the server owner also should receive pay - not only the internet provider. So, in with this idea, the server owner can potentially increase the amount of computing power to satisfy all requests.
- klez 1mo ago> there should be some cost for sending traffic somewher So now, because of bad actors, I need to pay for the privilege of watching a website or using a service that was meant to be free? No, thanks. I don't have a solution, but "break how the web currently works" is not one I would accept all willy nilly. EDIT: yes, I realize we already broke the web (with Anubis, cloudflare, recaptcha etc) but I think we should resist slowly breaking it further.
- feelamee 1mo ago> So now, because of bad actors, I need to pay for the privilege of watching a website or using a service that was meant to be free? No, thanks. First of all - I suppose it should be very cheap. So, real humans will not pay much. Second - why do u think that websites are meant to be free? They provide some service, so its a rather strange that the internet is so free (in both senses). I think, this freeiness is allowed to greatly speed up popularization. But for me is obvious that it can demand payment for service. And third - service owner really meant it to be free, I don't see any problems with this in my idea. It can still provide free service.
- klez 1mo agoI'm not saying websites can't demand payments for service, I'm just saying it's bad if it's a necessary fix for "scrapers are destroying the basic social contract of the web".
- inigyou 1mo agoThe internet is already like that, but for some reason the payment only extends as far as the recipient's ISP, not the actual recipient. Most senders pay a flat rate, but their ISP doesn't.
- delichon 1mo ago> Why is git.kernel.org “interesting” to crawlers Interesting to crawlers is not a narrow scope. We have the same problem on a B2B car wash site.
- iririririr 1mo agoso true. the article authors wishing crawlers will use git instead is so funny because the crawlers don't care at all. they are scrapping everything with brute force. they don't care about your content or effective alternatives, and one more site driving their real users crazy with Anubis is nothing more than a new blip in their dashboard. the crawler operators will not even look at the url.
- wiredfool 1mo agoSeeing the exact same thing on (somewhat high profile) open data sites I run. The crawlers get stuck in a loop requesting the dataset listing page with every. single. combination. of. facets. At essentially as fast as it can be pumped out or blocked. Had one bot super interested in a single organization to the tune of 1000 r/min for 24+ hours. Realistic rotating user agents, realistic sec-* headers, no ip address seen more than a couple times in 10 minutes. The only commonality was the route.
- marginalia_nu 1mo agoYeah my search engine saw traffic of up to 160 queries per second the other day from some bot that was ostensibly searching for information on Jack Parsons. Just variations on the same query in different permutations of filters and site:-terms.
- inigyou 1mo agoThat would be a different bot, one written specifically for your site. Mainly we're discussing the dumb ones that just crawl all possible http links
- marginalia_nu 1mo ago
- edent 1mo agoWordPress powered a huge number of websites. Yet the crawlers all go straight for the HTML of those sites rather than the more efficient and structured JSON API which all WordPress sites have. If these crawlers are so smart, why aren't they following the rel="alternate" which is provided explicitly for them?
- kardos 1mo agoBecause they suspect that, sometimes, different content will be served by HTML vs alternate APIs
- tptacek 1mo agoTavis Ormandy called this, about Anubis, almost exactly a year ago: https://news.ycombinator.com/item?id=44962529 https://news.ycombinator.com/item?id=44962529 It never really cohered as a solution. High-powered scrapers are better equipped to handle proof-of-work challenges than end users. Proof of work makes sense for a password hash, where any one guess at a password provides zero marginal utility. But every request from a scraper is productive to the scraper.
- nneonneo 1mo agoI disagree. The kernel finds it effective - 66% of scrapers are turned away directly. The scraper problem now is fleets of residential proxy devices - often things like smart TVs, phones, and browsers with some “proxy SDK” installed as part of an app’s monetization scheme. They make a couple of requests to a site - just enough to fly under the radar - and move on to a different site. If each new site they hit forces them to solve a proof-of-work, that’s a meaningful dent in their scraping performance. Many of these boxes may not even have the spare CPU power to efficiently solve so many proofs of work - and anything that makes an owner notice their device is running slow is something that could meaningfully impede adoption of these SDKs, or force the operators to choose between minimizing performance impact or scraping more sites.
- Y_Y 1mo agoThe implication here is that the proxy fridge forwards the Anubis challenge to a dedicated rig controlled by the scraper who efficiently solves it and returns the answer.
- wongarsu 1mo agoThat's still a notable step up in completely and resource investment for the crawler See also how captchas continued being effective for years despite services like anti-captcha offering to solve them for you for a fifth of a cent each by farming the work out to India. It took advances in AI that made it viable to reliably solve them on-device to bring the end of the captcha
- nneonneo 1mo agoI wonder if one solution here could be to turn up Anubis difficulty if the first page hit is not one of the obvious entry-points to cgit. It could even have a little hint that says something to the effect of “go visit the home page if this is taking too long”. (Better not to ban them entirely, in case people really did click on some random link e.g. in a news story or mailing list message). Distributed scrapers are going to generally try and hit their assigned list of pages; it’s a bigger waste of time if they have to go off to visit other pages first to get the cookie challenge.
- nxndbebdb 1mo agoAlmost all of my visits to cgit instances are through direct deep links. Hard to imagine someone randomly browsing git listings
- nneonneo 1mo agoThe kernel folks likely have a good profile on what page people trigger Anubis on (i.e. what page people hit first). From that they could make heuristics about what pages are likely to be useful deep links. Keep in mind that Anubis will rarely inconvenience an actual user who visits the site often; it’s meant to keep out first-time scrapers trying to grab a few pages from their queue.
- colinsane 1mo ago> I wonder if one solution here could be to turn up Anubis difficulty if the first page hit is not one of the obvious entry-points to cgit. Anubis has a fairly capable "policy" system. you can place something like this in your policy.json: ``` { "bots": [ { "action": "WEIGH", "expression": "path.startsWith(\"/expensive/endpoint\")", "name": "scrutinize-expensive-endpoints", "weight": { "adjust": 20 } } ] } ``` another thing smaller sites benefit from -- where the load induced by crawlers tends to be bursty (e.g. as they discover new expensive endpoints to crawl) -- is to adjust the difficulty up/down to maintain a steady system load. ``` { "bots": [ { "action": "WEIGH", "expression": "load_15m <= 16.0", "name": "sustained-low-load", "weight": { "adjust": -10 } }, { "action": "WEIGH", "expression": "load_5m >= 24.0", "name": "intermittent-high-load", "weight": { "adjust": 10 } }, ] } ```
- bauerd 1mo agoThey're the exception, not the rule. They get crawled like any other site, but happen to host git repositories. It's not obvious that these are targeted crawls and they likely may just end up in crawling queues a lot generally
- tarpitt 1mo agoMaybe you could have a system that heuristicially detects when an crawler is making the request and then feeds them a modified page, itself generated from an LLM, that injects vulnerabilities and bad code and discussion and such.
- NooneAtAll3 1mo agoit's hard to separate spambot that only accesses 3-5 links per IP and a legit user. Changing content for legit user can be devastating
- inigyou 1mo agoPut a cookie wall in front. The bot will either load the cookie and have a persistent identifier, or not load the cookie and not get in
- jay_kyburz 1mo agowhy is this not the answer? then you can also rate limit each cookie as well.
- sgsjchs 1mo agoit'll load the cookie, make one request, move to a different ip, load the cookie, make one request, move to a different ip, ...
- jay_kyburz 1mo agoIf you are discovering urls you have to wait for a previous request to finish. The rate limit should work. Requests without a cookie wait 2 seconds. Request with cookies can only make human scale number of requests per second? (1?) If you have a thousands of IP addresses, and you know all the urls you want to request in advance, you can just request them all simultaneously I guess. The next more advanced version is that URLs are unique to your cookie. Users can't share urls anymore, but it might be a tradeoff worth making. Unique urls for each user. You could probably still make this work, if you share your url with another user, they get the page, but heavily rate limited like a regular no cookie request. (a cookie url mismatch gets the rate limited version of the page)
- bluedino 1mo ago> But no, let's in fact choose the stupidest possible way of doing it — by rendering everything as HTML commit by commit and then parsing it. I feel like I'm at work. We had some web crawler using Selenium to make queries and scrape the data instead of just downloading the whole file. Every day it seems like we have some people that know just enough to be dangerous creating things like that. And then of course it's our fault that things are slow, or we won't give them infinite system resources, etc
- nxndbebdb 1mo agoJust serve the raw commit and render on frontend. I really don't get why they are complaining, just be performant
- RussianBot9580 1mo agoAnd for a shallow clone you would serve... what?
- hdbsbs 1mo agoNo need for a shallow clone, just let the Frontend fetch the relevant objects from a static file server
- NooneAtAll3 1mo agofrom what I see there are 2 solutions: 1) ban TV-proxy-as-a-service - straight up go to every representative there is and start pushing and lobbying and everything to stop spammers from distributing over non-computer devices, especially legally 2) make old commits more expensive to access than new ones. Legit users are not going to access those much, so they can pay the time. I assume diverse (unpredictable?) difficulty can also make spam pulling harder
- voakbasda 1mo agoI would think that residential proxying would be illegal already, as it is a network intrusion. The trick is chasing down the offenders, proving their actions did harm, and getting them to pay. None of those steps are easy, even if there are laws to assist. Otherwise, spam would be a solved problem.
- alwa 1mo agoIs it still an intrusion if the user accepted the shrinkwrap TOS of an app that trades them “free TV” in exchange for allowing that app to operate a proxy (via an “app monetization” SDK) on their network?
- voakbasda 1mo agoYeah, no one actually agrees to all of the individual terms in EULAs. That’s the first sign that the law will be nearly useless to address any aspect of these problems. It is already one-sided, and that side is not a friend to the consumer or general public.
- inigyou 1mo agoIt isn't illegal, there is no law against "network intrusion" which is a term you just made up, and if it was a real term it probably wouldn't cover this. There are laws against things like "unauthorized access to a protected computer system".
- 1mo ago
- znnajdla 1mo agoPut a CDN in front and let them absorb the load? Seriously, this is static content, which is so cheap to serve it should be free.
- ninglor 1mo agoThis is not meaningfully static content. Look at the charts in TFA. There is a combinatorial explosion of distinct URLs which the crawlers can and do request.
- grep_it 1mo agoDid you read the article? “[…]because we can generate 1.2 METRIC BAJILLION valid URLs just for a single fork of linux.git.”
- jwilk 1mo agoFrom the HN guidelines <https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html>: > Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
- inigyou 1mo agoWhat if it's really obvious they didn't read the article?
- znnajdla 1mo agoI did read the article. It just didn't occur to me that their combinatorial explosion of diffs was scrape-able. To be honest that sounds like an scrapers tarpit / honeypot now, because there is no value in scraping trillions of diffs. Sounds like the issue could be fixed by putting the diffs in a frontend app, not scrapable by URL, only by clicking around the app.
- jopsen 1mo agoThey allow you to diff commits, which is an awesome feature. But if bots a crawling diffs between all possible commits it's crazy. CDN will do nothing, because it's new urls each time. You can maybe find a CDN provider that block bots.
- deleted 1mo ago[deleted]
- gib444 1mo agoRunning Firefox with Temporary Containers Plus makes challenges 10x more annoying :D (Each new tab is isolated, unless opening a link in a new tab. Same as Safari in private mode)
- 0xbadcafebee 1mo agoI'm assuming they haven't yet sent responses to the bots? Since AI is dumb, you can send errors that tell the bot to git clone rather than crawl. If it's vulnerable to prompt injection, it might listen and do the clone instead and stop trying to solve challenges. Barring that, I think the solution is to charge money for access. Require users to sign up to render HTML, and provide a form of payment (any form you want). The cost is, say, $0.1 per GB. Rate limit all requests to reduce CPU. For the average user this will cost a few cents. For the bots you'll cover your costs and have a rate limiter to keep your system from being overwhelmed. Or they can git clone for free with no limit.
- jdnier 1mo agoI really enjoyed the writing style in this article. And the bot progression from "alter user agent" to "change IP addresses" to providers having to ban whole subnets, whole ASNs, and realizing "proxy SDK monetization" is a thing mirrors threat actor progression from the time before LLMs.
- TZubiri 1mo agoIt goes beyond mirrors, it's just something criminals have been doing since forever, to abuse all websites.
- atq2119 1mo agoIt feels like this progression of increasingly drastic measures to circumvent the protections of a computer system ought to be enough to establish criminal intent and get some of the people running those crawlers into prison.
- inigyou 1mo agoIt does. It's literally a felony but for some reason not a single person has pressed charges.
- NavinF 1mo agoturns out the overlap between "people who can't configure their webserver to serve at wire speed" and "people who can get law enforcement to take them seriously" is the empty set
- NegativeLatency 1mo agoClaude loves doing this on GitHub repos too, I have line in my agents file to tell it to clone to tmp and look there.
- ynniv 1mo agothis is an increasingly common situation. it goes something like: - i have a free, niche resource - it becomes too popular - i make it more efficient - now it's really popular, and people are "abusing" it - let's make them proof-of-work - ... and proof-of-work harder - but now "legitimate" users can't use it - ??? the core problem is that the average person uses a mobile device where work is expensive, and the "attackers" use servers where work is cheap. if you require expensive proof-of-work, next comes a cheap-work-as-a-service where inefficient mobile devices pay small amounts of money to get efficient servers to complete their work for them. now everyone has an interest in making their usage efficient, but there's still an obvious inefficiency in the system: why have people pay unknown 3rd parties to burn cpu cycles to reduce costs for a free service, when you could just have people make small payments that cover the service's costs? which is called l402/x402. micropayments' day has come
- deleted 1mo ago[deleted]
- inigyou 1mo agoWhat's actually stopping them isn't the PoW, it's the customisation effort. If one site has Anubis nothing scrapes it. If many sites have Anubis they write counter scrapers. Today if you make a slight change to the Anubis algorithm on your site, they'll burn CPU endlessly computing hashes with the original algorithm and submitting wrong ones. The author of Anubis hates this fact and will ban you if you mention it, so don't. He insists it's the PoW.
- ynniv 1mo agosounds right to me. ai will shred it
- Barbing 1mo agoSearching around for this, found "make users click the mouse three times"[1] as an anti-bot idea. Generally, makes sense that smaller site owners can make small customizations to existing anti-bot tech and see positive results until they're either (1) deemed valuable enough to receive custom attention or (2) the scrapers include LLM-based anti-antibot methods. >will ban you if you mention it Even if mentioned really politely? [1] would hope anyone trying this makes it accessible to visitors with disabilities
- Waterluvian 1mo ago> It was immediately extremely effective — the bots just gave up. For a few months, it was bliss: bots were blocked at the perimeter and gave up, moving on to easier targets; the users were mildly annoyed but tolerated it, and the Anubis stack was easy enough to deploy everywhere. As a tech person who works with tech people, I have become extremely sensitive to this kind of bias. Is this solution actually better? Is the CPU cost actually worse than mildly annoying everyone, or is it a problem being solved because it “offends the senses?” I’m not leaning towards yes or no for this instance. But I regularly see people jumping to conclusions without measuring. What is the cost of 20% and is that cost worth “mildly annoying” everyone?
- BowBun 1mo agoNot sure if you mean the solution, or the problem they were trying to solve. The problem is the cost here is paid by a group of volunteers, no? It's objectively reducing capacity for a really important public project by 20%. If you're responsible for keeping a public good like this available, it seems obvious to me to want to prevent this overuse of resources. Only one of the 3 parties involved here is not willing to engage in good faith right now and causing harm. I don't understand a need to try and tolerate them.
- inigyou 1mo agoBoth groups exist. Scraper DDoS is likely to burn 100% of your CPU on git diffs if you host git. But static file sites are unlikely to notice it.
- colinsane 1mo ago> Is the CPU cost actually worse than mildly annoying everyone > What is the cost of 20% and is that cost worth “mildly annoying” everyone? from the articled: > With a bunch of generous assumptions, legitimate requests are only about 2% of git.kernel.org traffic — everything else are scrapers. this is not some "CPU use is 20% higher than baseline" situation. it seems that people still do not understand the scale of these bad actors.
- Waterluvian 1mo ago
- singpolyma3 1mo agoWhy is no one filing lawsuits over this yet?
- Symbiote 1mo agoAgainst what person or entity?
- johneth 1mo agoBright Data et al.
- inigyou 1mo agoBright Data's business is legal, but they could be subpoenaed to find out which customer is making these requests, but first you would have to prove they were actually involved, because there are many residential proxy providers.
- singpolyma3 1mo agoWhoever is doing the abuse. If we don't know who that is we should find out
- inigyou 1mo agoSometimes the first step of a lawsuit is discovering who you're suing. It's not unusual and there are processes for it. You could bring something like an access log to a court and receive an order for all ISPs involved to unmask the corresponding users.
- mmooss 1mo agoIndeed. The solution isn't technical but legal. It's clearly abusive of - really stealing - other people's resources; there's no question about it. For some reason, like with fraud via email, text, and phone, we don't do anything about it. All this brazen crime and government does nothing; we don't even imagine government doing anything.
- 1mo ago
- iririririr 1mo agoanyone knows how Jwz solution is working? dont click next link because he will show a nutsack image if the referrer contains hackernews. love the guy. www.jwz.org/blog/2025/01/exterminate-all-rational-ai-scrapers/ basically, instead of blocking, he just poison it. and if a human sees it, it takes less effort to ignore the nonsense than it takes your pocket computer to deal with proof of work.
- semiquaver 1mo ago> because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge. This statement holds the core misapprehension behind Anubis. It’s not a ton of cycles. There is no difficulty setting that would be inconvenient for bots but usable for humans on mobile devices. I noticed the other day that lists.ffmpeg.org had moved to Anubis difficulty level 6, which takes ~180sec for my iPhone 17 to solve at ~100KH/s, making the site unusable. So I spent ~10 minutes vibe coding a safari extension with a native bridge to an optimized C kernel using ARM SHA256H* instructions that can do 200+ MH/s on the same device. This solves Anubis difficulty level 6 in a handful of milliseconds. Given the numbers and capabilities involved (a single $5K ASIC miner yields 200TH/s, a million times more hash rate than my optimized kernel running on an iPhone), I don’t see how proof of work could possibly be a sustainable strategy to keep bots out without ruining human user experience. It’s an arms race that can’t be won. Edit: I encourage you to try this yourself. Here's a sample prompt that ought to one-shot the task: > Build an iOS Safari Web Extension that accelerates Anubis proof-of-work using a native C ARM64 SHA-256 kernel. Precompute the invariant 128-byte challenge prefix, search fixed-width decimal nonces with ARM SHA-2 intrinsics and two worker threads, and target difficulty-6 solves under one second. Relay challenges from a Safari content script through the background service worker to native code, then submit the valid nonce/hash through Anubis’s normal pass-challenge endpoint. Include a deterministic benchmark app, correctness tests against CryptoKit, bounded execution, and fallback to Anubis’s stock solver.
- bjoli 1mo agoDoes it really take 3 minutes on your iPhone? My pixel 8 does it in slightly less than a minute in Firefox.
- efficax 1mo agoiphone 16 and i gave up after about 30 seconds, it was maybe a quarter done
- smallerize 1mo agoBut the scraper is making way more requests and is paying for all that compute.
- adverbly 1mo agoIs it really stupid if it means more data centers need to be built and it keeps the AI bubble going and GDP number go up?
- TZubiri 1mo agoSame problem we've been having for ages. Using shared ip banlists is the best solution so far, like cloudflare. Sure maybe they hit your server for 5 seconds and then desist, but they'll attack someone else, and they'll eventually rotate. I'm not sure if Anubis has a feature for centralized banlists, but I'm assuming since it's OS and privacy oriented, there isn't. There's a tradeoff between privacy and abuse, you want privacy? You get abuse, you want to battle abuse? Gotta sacrifice privacy. Worth noting that unmarked vpn users (residential proxy or residential vpn users) use these proxies for privacy, and therefore give a reasonable alibi to abusers.
- inigyou 1mo agoCloudflare doesn't block bots.
- TZubiri 1mo agoThat's the whole raison d'etre for CloudFlare, it was originally a DDoS protection layer, which, as the article mentions, is the final form of malicious traffic, being distributed and hard to attribute traffic to an identity. If you know CloudFlare as anything else, it speaks to how successfully it has grown and marketed itself into other areas.
- inigyou 1mo agoThat's what their marketing tells you it does - not what it actually does.
- marginalia_nu 1mo agoFWIW, git hosts have always interacted very poorly with crawlers, to the point where you have to actively code in git host detection to avoid getting stuck in an accidental crawler trap if you want to run a well behaved crawler. Easiest is just to look for anything that looks like a commit hash in a path and drop those URLs from the crawl frontier. Reason they interact so poorly is that is that git hosts generate a lot of links. One for each file in each commit, and a diff for each file appearing in a pair of commits. Even a small repo can have millions of viable links, and most of these are stupidly expensive to render for the git host. On top of this crawlers generally don't have a very deep understanding of what they are crawling, and can't meaningfully distinguish computationally expensive requests from cheap ones.
- Velocifyer 1mo agoI would add cloudflare, but set it to cache only mode *without* the bot blocking features.
- inigyou 1mo agoNo point, they are all unique requests.
- wingworks 1mo agoI think he means, get cloudflare to cache your content, so the traffic never reaches your servers to begin with. Assuming your sites content is cacheable by cloudflare. I agree it's a sad state of affairs if you have to rely on a 3rd party.. Maybe I'm not understanding how many requests at a time bots are sending to kernel.org (or how larger kernel is), but couldn't they have a local cache system too, where all it has to do it serve up dumb .html pages, needing next to no compute cycles.
- inigyou 1mo agoMaybe you're a little hard of hearing. NO POINT CACHING, THEY ARE ALL UNIQUE REQUESTS.
- tliltocatl 1mo agoIt is server-rendered cgit pages, there are potentially quadrillions of unique. They are not cacheable.
- andruby 1mo agoA creepy crawly is a South African invention to clean your swimming pool. The company that introduced them in the 70ies is called Kreepy Krauly. Also popular in Australia. https://kreepykrauly.co.za/about-us/ https://kreepykrauly.co.za/about-us/
- Symbiote 1mo agoIt's a childish word for an insect.
- leoqa 1mo agoIt seems clear to me we are moving towards a world where you will have to perform device attestation to access the internet. The spam/abuse is too great and accelerating.
- okanat 1mo agoWhat prevents people from obtaining or buying such devices and automating them? Using TVs as proxies is just one example of that. People will be willing to give their ID cards away too, if you pay them or beat them enough.
- inigyou 1mo agoIn fact they already do this. Buying 100 android phones and chargers is cheaper than reverse engineering whatever you're trying to automate - or was, before AI.
- leoqa 1mo ago.. and we ban those device ids and move on. Your capital is lost.
- inigyou 1mo agoEvidence shows otherwise. There are people making lots of money from these device farms. Right now. This isn't hypothetical.
- leoqa 1mo agoI literally work on this. The problem is that enforcing device attestation impacts DAU so platforms don’t want to do it but it’s effective. Once the internet is unusable, device attestation will be required to access services. Ad driven businesses will start to see demand from ad buyers for non-bot traffic etc.
- DarmokTanagra 1mo agoAI has simultaneously made the easiest parts of web development even easier while making the hardest parts near impossible.
- api 1mo agoThe AI companies should have their AI fix their crappy inefficient crawler code.
- jopsen 1mo agoI've seen this too. I think it's a few bad actors really. Because nobody serious about indexing content will do what these crawlers are doing.. They are consume lots of content that is unoriginal or duplicate or duplicate with minor modifications. Not sure how to block, but maybe a little bit of law enforcement could dramatically reduce the number of TVs being used a proxies.
- iamniels 1mo agoI run a website with 10k unique pages. If I leave the gates open, Meta hits it 200.000 times per day. Every day. What are you paying developers $500k for Mark?
- lxgr 1mo ago> [...] when a source is guaranteed to be LLM-free, like the entire history of kernel commits [...] Is that really the case? It was my understanding that LLM-based agents were explicitly allowed as long as their users follow certain guidelines [1]? And more generally: Somehow the theory of "essentially all bot traffic is AI labs crawling the Internet for LLM training data" doesn't make sense to me at all. There are at best dozens of labs capable of running their own crawl at Internet scale, but hundreds of millions of people using LLMs to answer their questions. (If my personal LLM usage is any indication, firing off dozens or hundreds of web fetches to answer a single question is not unusual.) While I understand that many existing projects have been resourced only for human readers and might as a result be struggling due to this, this characterization sounds a bit dishonest to me. And unfortunately, for this use case (i.e. ephemeral queries in a context possibly lacking storage or git access), forking the individual repo to answer a handful of string match queries against it might just be more expensive than to run that query against a web search index and then just fetch those results over HTTP. The solution would accordingly also look very different, as caching at the inference layer is significantly harder than at the training one (where it's most likely already widely done as that seems like a no-brainer). [1] https://docs.kernel.org/process/coding-assistants.html https://docs.kernel.org/process/coding-assistants.html
- desterothx 1mo agoThe use of LLM agents was allowed relatively recently, you still have 20 years of guaranteed no llm commits in the history. As for personal LLMs being the cause of the traffic, I think you are overestimating the number of users of LLMs generally, and even then most of those users are not actually developers. And there are even fewer cases still where the llm being used actually needed something from the linux kernel, personally after quite heavy use of LLMs on linux, I have never seen it try to fetch the linux kernel directly once.
- afarah1 1mo agoWhy not aggressively rate limit? Legitimate use of HTML rendered commits should be largely unaffected, and crawlers slowed to a halt. You can even jail after a number of 429's...
- mattmcal 1mo agoThere is a section in the article answering your question if you read it. > Suddenly, the crawlers were coming from millions of random residential or mobile IPs, all pretending to be random modern browsers. An IP like that would make 4-5 requests and then never show up in the logs again. There was no point in banning them, because by the time you figured out that they were bots, they were already done with you.
- mbirth 1mo agoI’d love a service like spamcop.net where I could submit my access_log and they lookup the abuse addresses and file abuse reports in my name. Maybe if people’s Internet access gets suspended they’ll think about installing random apps that work as a proxy in the background.
- inigyou 1mo agoabuseipdb.com Some ISPs ban customers based on a single report there - have fun!
- mbirth 1mo agoMany thanks! I've just submitted the first batch of 3000 (daily limit) IP addresses.
- ButlerianJihad 1mo agoThat is a ridiculous way to try and deal with the problem of residential proxies. You are, in reality, only hurting the actual owners, the subscribers of those ISPs who are behind those addresses. We call that "collateral damage". If any of those actual residential users try to use a website, their ability to freely access the Internet may be harmed by a bad reputation that they do not deserve. They may be totally unaware and non-consenting to residential proxy use. You are not, in fact, hurting the residential proxy-ers at all. Not one bit. They will move on to another IP and another compromised LAN, and they will continue to move on and on and on. They will not be harmed or impeded; they will simply keep turning up fresh, new, high-reputation IPv4 and IPv6 sources. This is a sheer numbers game, where the numbers are always in favor of the attackers. Also if network admins keep blocking/filtering abusive residential proxies, they will balloon their firewall rules and cause actual performance issues at the network level. You will turn into your own DDOS without any actual benefit. You're on the losing side of the numbers game, and in the immortal words of W.O.P.R., "The Only Winning Move Is... Not to Play."
- monegator 1mo agoThe thing that bothers me is why the fuck are they still scraping git.kernel.org or any other site that has already been scraped a million times before. Who would pay for that data? Then again there is the conspiracy theory about cloudflare sponsoring the scrapers
- alkonaut 1mo agoProof-of-humanity can’t come soon enough. We’re talking about privacy-preserving proof of age, but as we see here the real utility of such a system will be proof of humanity.
- Arubis 1mo agoHow do you define humanity? How do you ensure it includes every human? How do you ensure it doesn’t include every non-human? I’m not even asking about computation or algorithms. I straight up don’t think you can make a definition that isn’t a tautology or an approximation. Both of which are useful, but neither of which can fit a _proof_.
- Kamq 1mo agoYou're not wrong, but something can work well enough to still be useful despite falling short of the idea of a proof or any formal definition. Let's say that 95% of individual humans can pass it and only 2% of bots. For someone maintaining a website, who has to decide between using this system and shutting down their site because of the increased costs, that may very well be good enough
- Arubis 1mo agoThat’s a pragmatic and understandable argument. And for an individual hobbyist site owner, that’s fine. Are we okay with excluding 1 person in 20 from the services of a midsized organization? What if they’re integral to the workplace? Or a major transport provider without differentiated competitors? What if the organization is a state government?
- inigyou 1mo agoI'm one of the 5% apparently, cloudflare thinks I'm a bot. I suppose you'll exempt me in exchange for all my ID documents and bank statements?
- alkonaut 1mo agoEvery place on earth has some legal definition of who is human. The system I’m thinking of isn’t a technical/captcha one, it’s a human curated list of humans. Just an electronic ID. Those already exist but the challenge is making them (acceptably) privacy-preserving. I want to take my existing national digital ID and use it online basically. BUT I don’t want the websites to know it’s me. Just that I’m human (or perhaps over a certain age). And I don’t want the ID issuer to know what site/service asked whether I’m a human or I’m 18 etc.
- ChocolateGod 1mo ago> phone gets uncomfortably warm as it's doing the number crunching IMHO if I visit your website and it intentionally starts wasting my electricity for no other reason than to cost me money, with no opt in, it's hostile and malicious.
- inigyou 1mo agoYou try hosting gitea in 2026. The only other option is taking the site offline.
- stickfigure 1mo agoHow about just stop offering a html interface to the code? This doesn't seem like a critical service. Let people clone the repo normally. If someone else wants to run a public HTML service, let them deal with the bots. If you really want to offer a web interface, put it behind login. You can apply enough restrictions (captcha, super slow rate limit for new accounts) that it isn't cost effective to generate zillions of logins, and you can monitor logins for bot behavior. Sucks, but here we are.
- oasisbob 1mo agoI think you underestimate the difficulty in effectively gating user registration. If the anti-bot efforts in general don't work for other site pages, they won't work for signup functionality either.
- stickfigure 1mo agoI've been on the other side of this kind of thing (hero rather than villain, though I'm sure someone out there disagrees). You can fairly arbitrarily increase the difficulty of user registration, far beyond what users will tolerate for viewing an individual page. You can exploit this; a bot needs to make many accounts for the activity desired, and it's not hard to make "generate account" more expensive than it's worth for the amount of activity they get from each account. If it costs your attacker a penny to solve the captcha to make an account, and they can only get 100 pages out of an account, you win.
- TowerTall 1mo agoWhat if you make MFA mandatory. Normal users should have little issues with that despite the increased friction but bots would struggle with this.
- oasisbob 1mo agoI think this asymmetry is a fantasy. For many sites, your hypothetical penny to create an account is a cost an attacker would gladly pay.
- robotmay 1mo agoI've spent the last few days adding traps to one of my websites, ironically using LLMs of course, and I've been having quite a lot of fun doing it. Instead of the proof-of-work system of Anubis, I've gone down the iocaine route but implemented it in my application itself, as it's built in Elixir and causing problems for scrapers is really fun when it takes almost no server resources. Currently I trick bad scrapers into a fake infinite black hole path with the promise of tasty data, then serve images to them one byte at a time over 15 minutes (after sending the header quickly), bloat the responses to cost them tokens, and randomly return AI generated images of sexy toasters. I have an admin dashboard with a little leaderboard for which ones get the most stuffed, and it keeps my heart warm on these wet autumn evenings.
- someothherguyy 1mo agowouldn't that just make your connection load worse?
- robotmay 1mo agoCurrently not really an issue on this site but connections aren't an issue on Elixir usually anyway, unless you get up to about 1 million on one machine IIRC.
- someothherguyy 1mo agonormally depends on the type of traffic you are serving, but that seems like a weird thing to say as a blanket statement.
- jopsen 1mo agoProbably they must be deduplicating text they've seen before. The only punishment would be unique text that trains their models to be degenerate. And even then you'd probably have to serve across many domains.
- dspillett 1mo ago
- bilater 1mo agoInstead of trying to block why not monetize? So the proof of work can be directed at something you can be paid for (bitcoin mining)?
- inigyou 1mo agoMonero would work better.
- desterothx 1mo agoIf everyone had devices that could do some compute worth paying for, people would be doing the work all of the time. The problem is actually useful POW is a lot more expensive compute-wise, making it non viable for normal users
- pbronez 1mo ago“Expect to lose some functionality, at least when accessing our resources anonymously.” This seems fine to me. It would be a better world if we could have anonymous bulk data access. But if aggressive scrapers are bloating host costs, I’m fine with logging in. Now, the flip side is that ONCE logged in, I want my bulk access. The worst of all worlds with when you demand authentication and then STILL block bulk access. Case in point, I want to automatically download my Amazon and Target order records. This is easy to automate with playwright or whatever, but authentication stays annoying. My sessions expire quickly and I have to re-auth all the time. There should be an API to pull this data down.
- Kuinox 1mo ago1.4 billions requests, 258 160 cpu hours. That's 1.5 requests per second ? I'm starting to believe, the issue is more that their software is not well optimized.
- Skunkleton 1mo agoYou are forgetting that there are multiple multi-core servers.
- Kuinox 1mo agoIt's incredibly slow for a single core, that's my reference point.
- 6d6b73 1mo agoAdd a lot of random text to the html pages, preferably hidden to regular users, have the bots use lots of tokens to process it all.
- hnisjafx40 1mo agoLearned this the expensive way
- inigyou 1mo agoI also had this problem, but since nobody actually uses my gitea site besides crawlers, I just let a script run through my access log and ban every IP address who asked for a commit in the last 24 hours
- __MatrixMan__ 1mo agoThis is a fundamental flaw in the web. Since we treat a server's name as authoritative, anybody maintaining a replica needs to repeatedly hit that server to know if their remote version is up to date. If we trusted digital signatures on content instead of server names, we could have a model where a single bit of server load propagates to millions of interested parties. As it is the server must do something distinct for each interested party. CDNs mitigate this only partially, because mutable data means they have no good cache invalidation strategy. There's got to be a solution that doesn't involve heaping even more burdensome requirements on those who would dare to publish.
- dennis-tra 1mo agoContent-addressing decouples the hoster from the data itself. Anyone can serve content-addressed data and you can locally verify that you’ve got served the correct bytes. IPFS implements building blocks for such an alternative web. What irony that this article is about crawling content-addressed data.
- __MatrixMan__ 1mo agoAgreed on both points. But it's looking increasingly likely that particular dream is not coming true. Kubo, the reference implementation IPFS node, is maintainerless as of last week. I'm not sure where IPFS went wrong, but I think our web will continue to degrade until we figure it out. (It'll continue to degrade after that also, but then we can let it burn since we'll have a replacement to switch to).
- __MatrixMan__ 1mo agoI think focusing on filecoin was probably the mistake. You've got to build something that people trust first and then consider adding a money-shaped app. If you start with something money-shaped you're indistinguishable from the legions of scams, and that's a hard position to start from if you're wanting to build something trustworthy.
- inigyou 1mo ago
- mzajc 1mo ago> Why is git.kernel.org “interesting” to crawlers I think the post underestimates just how little thought and effort is put into these bots. I also run a cgit instance with far less interesting projects, and am not spared from the deluge of HTTP requests. The explanation I could come up with is that they try to crawl all links regardless of how much sense it makes or how much load it causes. cgit being cgit, this means billions of links for all combinations of parameters and hashes. That, or it's a deliberate DDoS attack.
- vintermann 1mo agoA lot of work is apparently put into bypassing any kind of anti-scraping, no work is apparently put into figuring out if the site freely gives a way to get all that information in a less wasteful way.
- sigbottle 1mo agoCode (not AI) in general is the worst at semantics. Not an excuse for their practices, but it makes sense they wouldn't try to figure out something more advanced. If they did, their solution would probably be something like, "try all current approaches we have"; it wouldn't be fine grained or truly reasoning at all unless they stuck an actual AI in front of it.
- TonyTrapp 1mo agoExactly my observation as well. They devour absolutely everything, no exceptions. No matter how stupid it might be to digest a source code repository via HTTP. They probably don't even recognize what's inside those pages and that there's an easier way to obtain the same result.
- arlattimore 1mo agoIn the case of kernel.org, why not make the unauthenticated version return only the latest kernel repo with no history (tiny number of URLs relatively speaking). If you want full kernel.org features, login.
- kuschkufan 1mo agobecause they do not want to be twitter, reddit, facebook, ...
- arlattimore 1mo agoI'm not suggesting 'login' because facebook/twitter/etc, just as a mechanism to make the bot problem go away. They clearly want it to stop, they tried obvious methods but the AI platforms are circumventing it (deliberately) which is poor form.
- kuschkufan 1mo agologin or you will not get the full service is exactly what facebook and co are doing and what your suggested would amount to. the kernel.org stated quite clearly they want to remain a public service. which i applaud.
- asah 1mo agoJust slow unauthenticated traffic to non-essential stuff...
- virgoerns 1mo agoI also run a public cgit instance and get over 1M hits every day, although my pet projects are nowhere near the size or impact of kernel. I had to block (via nginx conf) cgit endpoints for diffs, blame, snapshots and historical commits, because nothing else works. Now they return 402 (payment required). I consider this my total defeat and it's killing me inside, but it is what it is.
- inigyou 1mo agoYou could also publish a list of IP addresses.
- mzajc 1mo agoAs the article describes, it doesn't help, because the traffic originates from millions of unique residential IPs across hundreds of ASNs and countries.
- inigyou 1mo agoSo?
- VladVladikoff 1mo agoHave you tried blocking a million IPs before? Fail2ban gets pretty shaky at even 200,000 The AI crawler traffic I’ve seen sends one request per ip and seemingly has an infinite pool of residential IPs. You can’t block the ASNs becuase you also block honest clients. IP blocks are the wrong solution. And because I’m being negative I’ll also be constructive, IMHO the correct solution for fighting residential proxy crawlers is using RTT diffs this is one example https://github.com/Sakura-sx/Aroma https://github.com/Sakura-sx/Aroma
- pmlnr 1mo agoFail2ban becomes a serious bottleneck at significant traffic. I've replaced it with a shell script and direct pf commands that run every few minutes.
- superjan 1mo agoHow feasible would it be to only offer a binary git (partial) download and move the html rendering to the client? It would still be a lot of requests, but less work for those servers. Not that I like SPA’s, but they could be useful here.
- forrestthewoods 1mo ago> Shallow clones are awful. Run your own damn mirror if you're going to do something nasty like that. TIL shallow clones are expensive. That's wild to me. It's supposed to be cheaper!
- Backslasher 1mo agoI thought they were expensive compared to fetches from established repos. TIL they're also expensive compared to full clones.
- cobbzilla 1mo agoI ended public access to my git server after I got flooded by bots and my own commits were noticeably lagging. That’s not an option for the kernel. It’s hard to read the cat-and-mouse account with any hope today. I think the flood abates someday but not sure how it happens.
- golem14 1mo agowhy not require an account for html access and otherwise just serve say a pre-cached complete repository requiring minimal work?
- bourse_lee 1mo agoWhat if Anubis computations were turned into a crypto-miner
- fer 1mo agoBrowsers, at least Firefox, blocks crypto-miners. How to tell legitimate from underhanded crypto-mining?
- throwawayffffas 1mo agoThe whole concept behind Anubis is flawed. It tries to block access by imposing a compute cost to block people that are using incredible amounts of compute just to generate and parse the requests. The people that Anubis tries to block have tons of compute to spare. For example at runpod, if you rent a container with just one B300, you get 32 cpu cores and 250 gbs or ram that are essentially just sitting there while the gpu does all the work. If you think you can impose a compute cost on these people without blocking everyone else, well think again. Or in my case, I am running qwen3.8 at home on a couple of gpus, these are attached on 32 core epyc server with 128gb of ram I am pretty sure i have more compute than the typical dev laptop. I am of course nice, and don't aggressively scrape peoples services.
- calvinmorrison 1mo agoI love to see the 'leet kernel hackers and maintainers' struggling with basic volume. Each page load should cost you near nil. Us lowly PHP developers have been caching shit for close to twenty years. Learn how to cache your application and your cpu usage should be almost zero. In fact basically any read should cost nothing in comparison to writes.
- inigyou 1mo agoIt's running a diff between two arbitrary blobs of text. Do you actually have a solution or are you just saying to remove the feature from the site entirely? > Us PHP developers I can tell.
- calvinmorrison 1mo agoit's not an insult. yes. cache heavily. shitty php stacks serving trillions of dollars of ecommerce sales have managed to do it for a long time.
- zbentley 1mo agoThose ecommerce stacks serve a large fixed number of pages. cgit does not. Imagine if WooCommerce had a route "/product/<sku1>/compare/<sku2>" which displayed an auto-generated comparison between any two product pages. Now imagine running a million-SKU WooCommerce site, where each product page was 100kb of text. Now imagine scrapers are permuting those URLs. How would you cache that? That's what cgit/kernel.org and many other Git forges are dealing with. These aren't static websites, even if the underlying Git repo is largely static; they're rendering arbitrary diffs and other generated-on-the-fly views into Git history. The ability to do that is a large part of the value of a Git UI.
- inigyou 1mo agoCan you explain how a cache lets you avoid serving a request for the first time?
- riknos314 1mo ago
- hei-lima 1mo agoGreat chart! Does anyone know what tool was used to make this?
- userbinator 1mo agoSo, you'd think that something that pretends to be “Artificial Intelligence” would use the most efficient way of using our data for training purposes, right? I'm strongly convinced that these aren't "AI crawlers"; they're just plain DDoSes done by those who have interests in turning the Internet into a dystopian walled garden with "security", and now they have a convenient scapegoat to blame. Don't you find it too coincidental with the rise in identity/age verification and other attempts at silencing free speech on the Internet? It's widely known there are questions that LLMs can't solve, and once in a while an obvious example appears, so a simple CAPTCHA-like challenge with an HTML-only form would be the logical "defense". Instead there's a huge interest in pushing JS-required proof-of-work (as others have pointed out, these attackers have far more compute than the average user) and remote attestation (there are already providers with huge farms of mobile devices that can defeat this easily). Things just don't add up.
- desterothx 1mo agoThe questions the LLMs can't solve puts you in a cat and mouse game. They sure can solve most of those with the correct tooling, and i would argue its easier to create this tooling than it is to keep innovating with new questions LLMs cannot answer
- userbinator 1mo agoThis is already a "cat and mouse game", but one that conveniently benefits both the browser monopoly and the authoritarians wanting to turn the Internet into a closed proprietary system. ...and how convenient that pointing out the inconsistencies and this massive scheme of propaganda and manufactured consent gets you downvoted. "Truth does not resist questioning."
- kgeist 1mo agoHow about: "Type the seahorse emoji to solve the CAPTCHA" :) Something that triggers infinite loops in LLMs or trips the guardrails.
- RGamma 1mo agoKitboga (guy who trolls scammers) has some funny CAPTCHA setups if you need inspiration. E.g. https://youtube.com/watch?v=TOzEnwl7LkA https://youtube.com/watch?v=TOzEnwl7LkA
- oowa 1mo agohave a hackathon to solve for this. OP says it's not a problem for him right now but if we extrapolate what he's talking about it's definitely a problem aaaaaaand It's totally solvable, Even with all of the crazy combinations he's talking about it's still solvable. And it's already been solved using patterns we see in streaming services. This is completely hackathonable. but why do we even need to bother with this? The slurpers are the cause of this, and they can cause this problem because of Murphy's Law. well you can only account for Murphy's Law with good architecture or something like that or whatever. Ha ha hackathon.
- stefantalpalaru 1mo ago[dead]
- javcasas 1mo agoAt this point they are using residential proxies and stuff, and increasing the difficulty level is not going to help, among other things because they don't pay for it. Why cannot we turn this whole proof of work thing into an official "help mining $SHITCOIN"? I mean, if they really want the data that badly, at least have them pay the hosting with their CPU/GPU/ASIC cycles.
- beached_whale 1mo agoI wonder if they could pre-render the stuff older than a month ago and compress it and serve it as static content. Not optimal, trades space for CPU, but might be cheaper.
- kdowns 1mo agoI made it to a third round interview at anthropic in 2024 and they had me build a web crawler as their programming test. Part way through I started on making it respect robots.txt and I could immediately tell they were no longer interested in me.
- myng111 1mo agoThat's really funny. It's always been kind of amusing to me that Anthropic has this air about them of trying to be the most ethical AI company, but really exhibits the same behaviour as all the others.
- duplessitous 1mo ago> Training an LLM on content produced by the LLM gives it the equivalent of a digital prion disease, so when a source is guaranteed to be LLM-free, like the entire history of kernel commits, it's worth its weight in gold as a source of training data. Except the majority of LLM training content nowadays is synthetically generated by LLMs. I wish people would stop making this statement, I don't know why this claim persists to this day. It wasn't true two years ago and it sure isn't true now
- thomasjudge 1mo agoAre there lots of people doing development on mobile devices?
- goldenarm 1mo agoMany are on old laptops, which suffer the same way
- innocent_name 1mo agoWhy can't they just ask Linux Foundation for 96, or even 1696 cores? If you're reading this - go ahead and see HOW Linux Foundation spends their money.
- _blk 1mo agoWhy not use the POW to help cover the costs? Mine an actual coin (Annubis Coin?) and pay for anonymous infra access with it (or log in and get a certain quota for free)?
- ivanjermakov 1mo agoThere has to be some not-yet-discovered way to have a capcha that is easy for any human but impossible for robot. Too bad capchas hurt user experience no matter how easy they are. Another solution I came up with while reading HN comments: whitelist IPs instead blacklisting. Give access to well-behaving hosts/groups. It can even be shared across different sites. Although this would create a market for selling "good IP" proxies.
- brownkonas 1mo agoThe reverse turing test :(
- echelon 1mo agoEven better: make people pay for access or vouch to give access to a third party.
- ivanjermakov 1mo agoSome kind of proof of stake might be viable. "I as a visitor stake 1 cent that I'm a genuine user and not a sloppy bot, server is free to withdraw my stake if it's not true". If works, withdrawn money can be used to cover hosting costs.
- zythyx 1mo agoThere are definitely ways of proving you're a human, unfortunately it also means giving up your privacy and anonymity (IRL ID Checks combined with appropriate routing and validation - even going as far as certifying the browser being used) Obviously though none of us want to give that up, so the alternative is that we can almost never 'prove' we are human especially with bots getting as smart or smarter than the average redditor.
- Barbing 1mo agoImagine proving you have no financial incentive to get on the whitelist. Reminded of the mules renting Airbnbs to use as USA-based delivery locations (tricking grandma into FedExing cash for one scam or another) - https://getrichslowly.org/scambaiters https://getrichslowly.org/scambaiters (probably Jim Browning + Mark Rober specifically https://youtube.com/watch?v=Xvjjpzyiig4 https://youtube.com/watch?v=Xvjjpzyiig4 ) But! Using a network of real ID-checked humans to scrape the web, what would that be--half a billion times harder than Firecrawl or whatever they use today? Too bad it's dead in the water today because so many (like me) hate the idea so much. Perhaps a biometric dongle (retinal-scanning orb :-/ ) that the staunchest privacy hawks stamp with their seals of approval because it's somehow nearly impossible to go horribly horribly... horribly... wrong... Yeah, anybody who can crack this issue, hope you have the free time or find the funding to try it, we need ya.
- cute_boi 1mo agoAll this happens due to companies like browserbase, Hyperbrowser, Scrapefly. These service exists to facilitate such operation and they aren't doing anything to prevent abuse. They are infact selling way to bypass captchas etc... I think any service that is trying to sell a way to solve captchas must be banned by government. At least these things shouldn't be done so openly.
- fizlebit 1mo agoI wonder if we're back to peer to peer networks with proof of useful work (e.g. serving read requests) vs proof of wasted work.
- DrJThomasHusk 1mo agoWhen a hapless user visits my site well muahahahahah Sorry, just the thought of it But when they do… boy do I have a trap waiting for them. My wife calls me The Genius. I’m the guy she calls when her battery dies or when her instagram breaks like when it shows that random guy in her DMs, stupid bugs LOL I digress. Alas, when a user lands on my page. My page wants to know exactly 2 things: 1. Why are you here and who are you And 2. Can you produce a working solution to Pharoah’s Fortune …those of you aren’t familiar Pharoah’s Fortune is an old chestnut little poem, a riddle if you will I like to ask candidates and so far nobody’s solved it And the reason nobody has solved it is Pharoah’s Fortune is a very tricky problem. It’s not something you can “solve” per se it’s more like you arrive there. So far no one has solved it. They all fall for the same trick! It is of course what separates those who write elegant C versus those write poor quality JavaScript. So I always say to my students to keep an open mind because you never know who - or should I say where you’re talking to. I’m bookish.
- phyzome 1mo agoI don't see any relevant reference online to "Pharoah’s Fortune".
- boredatoms 1mo agoCan I suggest putting some text in the page that tells the bot what the more efficient download method is?
- zbentley 1mo agoWhat makes you think that would accomplish anything? These scrapers aren't LLM agents. They're distributed classical programs that harvest data which is later used to train an LLM. The LLM doesn't write the scraper or respond to individual scrape events. The entity training the LLM contracts someone, who contracts someone, who contracts someone to run a web scraper and send them the data.
- bjourne 1mo ago> They still do that — welcome to the wonderful world of “proxy SDK monetization.” It's big business, and your TV is probably doing it. I must be missing something. How can using peoples' TVs as bot farms be even remotely legal? Especially when the purpose is to avoid IP blocks?
- sgsjchs 1mo agoThe people "consented" to this when they clicked OK on the user agreement.
- yardstick 1mo agoHow about Allow git clone for free/unrestricted still. Require the user to sign in to view html views. Sign in require a valid email or phone where a validation link is sent. Or: Users signed in won’t see the Anubis. Users not signed in can still see the html views but have to use a very high work level? Or: Limit unauthenticated requests from an IP to 5/minute. Authenticated requests can do a lot more before hitting the limit.
- Barbing 1mo agoWould it help if this only applied to old pages? Want to render than seven-year old commit via HTML? Sign in. (But I have no idea.)
- talkingtab 1mo agoTime to F*$k the internet. The whole concept of anonymous IP addresses was broken but worked for a long time. Now it is just stupid. Just like domain names. (Are more names used by squatters than real?). And email as identity? Time to engineer solutions and create a new protocol layer. This is not a hard problem. It just requires that someone build a certification wall. The IETF should have done this long ago, right?
- desterothx 1mo agoIt's not a hard problem, the hard part is doing it while keeping a similar level of privacy/anonimity
- hamandcheese 1mo agoI wonder how much is for training vs for LLMs doing research. On several occasions Claude has gone digging through kernel archives on my behalf (sometimes at my direction, other times all on its own). Usually to determine the current status of some kernel bug I'm experiencing. Apologies for the load, but I'm sure it was much less than an actual crawler trying to slurp up everything.
- iLoveOncall 1mo ago> Where does that leave us? Honestly, the answer is simple: sue. It'd be hard to argue that it's not a DDOS.
- oowa 1mo agoTLDR basically old tech is not optimized for scrapers / slurpers / etc. to the point it would take 42^n to solve all possible combinations. Why? Murphy's law. Solution for OP is to ignore for now. Otherwise Use or invent something else. Easy enough. other notes... Anubis and other gatekeepers dont work perfectly, but ok for now.
- louiskottmann 1mo agoIsn't it perfectly reasonable to require an account for any use, and to ensure that making one has a high level difficulty anubis challenge or delay ?
- andrewaylett 1mo agoJust for fun, because I could, I vibed up a `cgit` replacement that runs entirely in the browser -- point it at a git repo where you've run `git update-server-info` and it'll load files as if it's starting to clone the repo, using range requests and browser caching to avoid actually loading more data than necessary for the view you've requested. I'm certainly not saying you should use this code, but it's a proof of concept for avoiding the CPU overhead of cgit rendering by loading the data on the client. It cost me £8.27 of Fable use (from the free credits I've been given) and 56% of my five hour quota on a $20/month Pro plan. There's no server logic, it's 1.3MB of minified JS and CSS and (while I'm absolutely not suggesting anyone try to use it) it basically works: https://github.com/andrewaylett/rgitweb https://github.com/andrewaylett/rgitweb This is a one-shot, my prompt set the expectation that I'd be able to load resources using CORS but (not entirely unreasonably) the Git hosts I've tried don't set CORS headers. Shared more because I was pleasantly surprised at how cheap and easy this was -- and with a repo link because talking about it without sharing the link would be a bit crass.
- andrewaylett 1mo agoNow deployed here: https://andrewaylett.github.io/rgitweb/ https://andrewaylett.github.io/rgitweb/ with a copy of its own repository to explore.
- CqtGLRGcukpy 1mo agoI just had a look through my logs, and I've had over 90,000 requests from known AI bots over the last month. All this to a personal website that doesn't post very often. And that's just known AI, I can't imagine what requests are pretending to a real person when they aren't.
- justAnotherHero 1mo agoWhile nowhere near compared to their scale, I run a consumer app where most of our users are using the mobile app, with the web app getting perhaps 10-15% of the mobile active users. However day after day it just gets blasted with requests for deep pages. I was quite alarmed when I saw a 100x increase in the daily active user numbers which relied on session length, only to realize they were all bots. Naively I too initially resorted to blocking user agents(Meta is thankfully nice enough to identify themselves, not nice enough to stop blasting 50k requests a day however), IP ranges from cloud providers and various browser fingerprints that I found connected to suspicious traffic. However the battle seems unwinnable at the moment, outside of gating all content behind auth which I don't want to do. We have around 500k user generated content pages and I want those to remain publicly available. I would be happy to provide our data to any one of these scrapers and I even added a message asking them to contact us if they want access to our data whenever I return a 403 response, however nobody has reached out. Another campaign that someone is constantly running is daily checks for 100s of possible secret/config paths in hopes of finding an exposed private variable, these i've just blocked even though they would return a 404. I still haven't found a way to deal with rotating residential IPs however, and most likely never will. My current approach is to just run a 24 hour scan of all requests with codex and update my next.js proxy with more IP ranges, browser fingerprints and anything else that won't affect a real person. Has anyone managed to come up with a way to stop this onslaught of crawlers and scrapers?
- timpera 1mo agoI had the same problem, tried a bunch of stuff, banned millions of IPs, but nothing really worked. I ended up spending a bunch of time rewriting the website with Codex so that it is more efficient. Now I'm still getting the same junk traffic, but it barely affects the CPU anymore. I think this might be the way to go.
- vist_orn 1mo agoReal creepy crawlies in the server rack are always a bigger surprise than any code bug.
- greatgib 1mo agoI have the feeling that the hate might be misplaced. For a shopping website or user generated content website, I might understand the terrible load of crawlers that are trying to "steal" the data. But for the kernel, what's the purpose? Are you that "no human" are seeing your page or its content? Maybe we should investigate more the usage being this "bots". I don't buy the explanation that there are millions LLM that are constantly trained on redownloaded data from kernel.org. What would be my better guess is that it is not training, but users are actually accessing this content through chatbot and co. Like when you ask why your sound is suddenly not working anymore after an update or why your wifi driver is constantly disconnected after leaving sleep, it might be possible that the "LLM agent" is requesting the commit contents to "understand" or refer or explain them. Is it a bad thing if it helps users? But actually, regarding this article, I'm quite amazed that with all the advances of the linux kernel, and server softwares, and that the C10k challenge is solved since a long time, still such a basic traffic is such an issue. > At any one time, across 5 geo-distributed nodes, there are 14 CPU cores doing nothing but rendering git commits as html. 14 cpu looks nothing to me. It's like you have 1 iphone and 1 raspberry pi active in a corner of a room. Counting in "seconds" of activities, easily shows meaningless huge numbers. Do you want to know how many breaths I take per year? 8 to 9 millions! Most certainly, the usage of this shitty Anubis has ruined the climate million times more only with the wasted cpu resources of legit users... But moreover, by definition the git commits are not supposed to change, ever, so can someone explain to me why the fuck do kernel.org "re-render" the commit to html each time someone is accessing it instead of using a cache or a static version of the html of this commit? > oh, several BILLION valid URLs you can scrape, only to get 922 duplicates of the same 1.48 million commits Again, reading that, my immediate thinking is that it is a shame that such talented people would not be able to have a proper optimization, so that getting the 922 duplicates are just costing a fraction millisecond more after the first person retrieve the first page.
- zbentley 1mo ago> the C10k challenge is solved since a long time This has nothing to do with that. Any Node.JS application will happily accept 100K connections. They'll all wait for the under-resourced database behind it. That application "solved" the C10K challenge, but it's still overwhelmed. > it might be possible that the "LLM agent" is requesting the commit contents to "understand" or refer or explain them. Is it a bad thing if it helps users? The article describes random algorithmically-generated traffic arriving in batched waves from laundered residential proxy IP addresses, a few unrelated hits in a group then gone. That's not the pattern you'd see if end users were asking their agents for help. > it is a shame that such talented people would not be able to have a proper optimization It's mostly not static content in the sense that you're implying. Routes that access a single commit can be cached. But most of the routes scrapers are hitting are e.g. computing diffs between arbitrary pairs of commits, or other computed-on-the-fly views into history. I'm sure they're already caching their useful-to-real-humans data. As the article said, the vast majority of their traffic is bots hitting those arbitrary, permuted URLs. So whatever cache they're using is probably a) missed almost every time, and b) constantly getting evicted to make room for data served to bots (unless they eschew caching to avoid this--fair--and are thus back to the original issue regardless). There is no "proper optimization" here. It's not slow to go compute the diff between a random pair of refs, render that into pretty HTML, and serve it. But it costs something more than a cache hit, and doing that dozens-to-hundreds of times a second constantly consumes resources.
- emsign 1mo agoWhat a waste of energy LLM training is. Meanwhile Himalayan mountains are crashing down. I love this world. It's so idiotic.
- sgsjchs 1mo agoIronically, defense by obscurity may be the way to go here. Fork Anubis. Slightly modify the hash function it computes. Deploy. Do not try to make your fork widely adopted. Do not even publish it. You've just defeated ASICs and any craweler that's special-cased Anubis (currently all of them). If enough people do this, the only recourse they will have is either genuinely executing served js code like a real user or building some unholy pipeline that uses ai agents to compile it to a GPU kernel for every host.
- slipknotfan 1mo ago> Fork Anubis. Slightly modify the hash function it computes. Deploy. Do not try to make your fork widely adopted. Do not even publish it. A more robust solution would be to keep a few patches handy with different versions of the algorithm, and rotate which one is in use. This would keep the crawlers on their toes if they wise up to the changed algorithm. One could even imagine automatically rotating witch algorithm to use on a weekly basis.
- chr15m 1mo agoIt doesn't matter what the hash is if it is inherently cheaper for a bot farm to compute the hashes than it is for a human to do it on their device. The human pays a greater cost in annoyance, wasted time, battery, and that means the PoW has failed its function. The bot farm owner does not care.
- afdbcreid 1mo agoIt's not really defense by obscurity (the JavaScript is public), more like defense by... being different?
- mattacular 1mo agoSnowflake defense
- rtpg 1mo agoI was thinking about this. Would it be possible for Anubis to have a code gen process that would make every deployment of itself sufficiently unique such that it's "annoying" to work around? I do wonder about the premise as well: are people special casing for Anubis?
- dunlin 1mo agoReminds me of debugging production issues at 3 AM. Both can make you jump out of your skin.
- lmz 1mo agoIt's funny how some people say "AI bad, datacenters waste energy" then other people say "AI bad, going to make humans and their phones waste energy".
- gruntled-worker 1mo agoThere's something missing from the picture. The bots are: - Using a terribly inefficient way to redownload the same commits as e.g. HTML diffs, possibly the most inefficient. - Putting in tons of CPU cycles to surpass the Anubis PoC. - Putting in other kinds of active effort like reworking access methods and buying "residential proxies" that are probably illegal in most jurisdictions. This sounds more like escalating DDoS than AI scraping.
- zbentley 1mo agoPossible, but I doubt it. The sheer number of other free-to-read content sites dealing with the exact same problem the last few years (many of whom there's no plausible reason to DDoS) tells us that this is content harvesting, not an attack.
- gruntled-worker 1mo agoI'm curious and would like to see those reports. There's AI scraping for sure, but intentionally resource-consuming, increasingly-insidious AI scraping I've never actually read about. AI scraping might be bad, but if a particular case that's actually a DDoS becomes the cause celebre against AI scraping, it will weaken the argument, not strengthen it.
- adangert 1mo agoCurious, if serving bots (and traffic) is the main concern here, why is a distributed git solution like radicle not considered? https://radicle.dev/ https://radicle.dev/
- ironqcold 1mo agoMy takeaway: we're degrading the web for real people to slow down bots that will just move forward. The solution seems is worse than the problem. At some point, we need to accept that the open web as we knew it is dying...
- kristianp 1mo agoOne problem with Anubis is that once you've solved the POW once, you just need to hold the cookie to avoid solving it again. Scrapers have probably learnt to do that by now. So Anubis isn't as effective as it used to be before it was widely used.
- UltraSane 1mo ago$1 dollar a year subscriptions would help.
- DarmokTanagra 1mo agoyou do realize crypto has faced this problem, presented the exact same solution, and utterly failed right?
- UltraSane 1mo agorequiring payment requires a payment method and burning those is much more painful for whoever is running a scraper.
- Wowfunhappy 1mo agoWho exactly is running all these scrapers? There are, what, maybe 15 major AI labs, if that? And none of them are smart enough to realize they could just `git clone` all the content and use it offline?
- strix_varius 1mo agoThere are many more labs than that, and humans aren't designing unique scraping processes per domain.
- micah_chatt 1mo agoIf you do the math (also a common system design interview question for an AI lab), its actually only ~3PB (compress to 1PB, ~$22,000/mo in S3) and a few thousand/mo in compute over less than 4 months to index the entire internet for pretraining purposes. At that price point, its actually very affordable to many thousands of organizations to get their own copy. I would expect the major labs to special case kernel.org similarly to other sites like Wikipedia, but not the majority of scrapers
- Wowfunhappy 1mo ago> its actually only ~3PB (compress to 1PB, ~$22,000/mo in S3) and a few thousand/mo in compute over less than 4 months to index the entire internet for pretraining purposes. That is super interesting, thank you! > At that price point, its actually very affordable to many thousands of organizations to get their own copy. I'm still confused as to who is actually doing it though! Maybe it's affordable to scrape and store, but training a competitive AI model is going to cost much more, right?
- DarmokTanagra 1mo agothree guesses what geographic regions they are based in.
- GoblinSlayer 1mo agoI suppose proxies provide only http request interface.
- CGamesPlay 1mo agoGit forges seem especially prone to this: tons of information, highly valuable to scrapers, rendered through several different lenses, gives a combinatoric explosion of URLs. Obviously scrapers could just be less stupid and clone the repo, but it's not happening. It feels like turning these frontends into JS-only viewers would resolve this, for the most part. The JS clones the repo in memory and renders whatever lens the requestor wants, and the server becomes a dumb object storage that uses less resources. The anti-JS folks are free to clone the repo still, and view whatever lens they want, so that minuscule slice of the legitimate requests is still served, albeit with a degraded experience.
- zbentley 1mo agoThat'd fail for three reasons. First, scrapers would start running JS. Whether they're running chromium-in-a-box or something more clever doesn't matter. Compute is effectively free for them, and even shitty WebOS set-top boxes can probably run a stripped-down headless browser with a JavaScript engine. Second, a full clone in browser memory is massive for something like the Linux kernel. Lots of browsers and pro-JS folks wouldn't have the resources to run that. Third, assuming that scrapers are willing to play the JS game, suddenly you have a massively increased rate of full clones happening. Even if serving raw git is cheap for your backend, the bandwidth the elevated clone count drives is not.
- CGamesPlay 1mo ago1. Preventing access to open-source code by scrapers is not the goal behind this idea. The goal is to reduce the server load taken by the scrapers. 2. Why would you need a full clone to access /blob/cee9395acd8043be0644b25c34bfa86623f2b935/block/badblocks.c?
- zbentley 1mo agoI don't think the show-commit or file-at-revision routes are what's causing the bot load. Rather, the article describes the routes that compute views (e.g. show the change history across several commits) into history as being the issue. If we take that example, a JS client would have to fetch N individual commits/routes/objects and compute the requested view, but first it'd have to fetch the indexes/logs to determine what commits exist within e.g. a specified range. I suspect that'd require more work on the frontend than "just run WASM-built git" ... unless the proposal is for it to fetch all requested objects lazily, in which case I think you'd be surprised how many files are read by git when answering a question like "show me the diff by user XYZ in file ABC on branch QRS between date 1 and date 2". That starts to get expensive to pull in the browser, and the bandwidth costs might start to hurt even if the backend now only had to serve cacheable dumb blobs.
- wolttam 1mo agoI think the solution is for the POW being done by the clients to *actually benefit the site owner*. Users remain just as mildly annoyed as with Anubis, but maybe a bit less knowing that the work they’re doing benefits the site owner/author, and the system helps thwart the bots (or at least makes them do work that benefits the author).
- dzhiurgis 1mo agoThese are most likely not training scrapers, but people looking for concrete pieces of information (i.e. commit, comment, etc).
- dzhiurgis 1mo agoClarification - these are crawlers that llm’s use when you ask them to check this site for foo.
- 0xdeadbeefbabe 1mo ago> permanently tying up a chunk of capacity spent on producing output that is only useful for a single purpose — feeding a learning model. The horror.
- BorisMelnik 1mo agoit's getting insane, I have a high profile client, I manage their infrastructure including web server. I swore to them years ago they would not have to turn on the CF managed challenge / under attack / human verification. I've handled every type of attack and malware that came their way but these past few years, ai scrapers are a large por or their traffic, eating into the budget and now interfering with sales. and I don't know if anyone else is noticing or watching these ASNs but it sure looks like a few well know and big name AI companies are using *residential proxies* to so their scraping.
- charcircuit 1mo agoHow about making cgit more efficient at serving these pages. There's no excuse for burning a ton of CPU power on purely static pages when you have generous resources available to you.
- stcg 1mo ago> proxy SDK monetization Wait what? I never heard of that. I call that a botnet
- GoblinSlayer 1mo agoIt's called capitalism.
- hubraumhugo 1mo agoThere is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race? Some approaches that I think are promising: - A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.). - Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely. - what else? [0] https://datatracker.ietf.org/doc/draft-vaughan-machine-readability/ https://datatracker.ietf.org/doc/draft-vaughan-machine-reada... [1] https://datatracker.ietf.org/doc/html/draft-meunier-http-message-signatures-directory-05 https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...
- sunaookami 1mo agoHave the exact same problem with my MediaWiki and Gitea instances. Doing a managed challenge through Cloudflare (sigh) on "expensive URLs" helped and minimizes the impact on real users. These are also all over a million of residential IPs, very annoying.
- runtime_lens 1mo ago[dead]
- yread 1mo agoWe need a solid way to prove we are meatbags. How about a simple USB accessory that goes around your neck and measures your ECG. It could also give you subtle electrical shocks (in a unique pattern; a challenge) and measure how the ECG responds to that.
- desterothx 1mo agoim rather fond of a blood tester that gives you unique drug cocktails as a fingerprint
- yuumei 1mo agoHere is an idea: instead of PoW do “proof of AI” each user who wants to access the page run AI/LLM prompts to undercut the corpos abusing the pages
- smetj 1mo agoOffer the content, which is otherwise scraped, as a tarbal + diffs via a torrent or p2p network making it more efficient and cheaper for the crawlers to obtain the data.
- deleted 1mo ago[deleted]
- aimen2 1mo agoNow hosting will cost more for everyone...
- akoboldfrying 1mo agoAre you prepared to randomly serve data that is incorrect -- but in such a way that real people can easily detect it -- some very small fraction of the time? If so, you could serve iocaine-style bogus pages 1% (say) of the time that: 1. "Look like" real pages to an LLM-less computer (if you get to the point where you have pushed crawlers to use LLMs to detect nonsense, that already increases the cost a lot) 2. Look "obviously wrong" to a human (E.g., you could take some regular text and swap the order of each adjacent pair of words) 3. Are cheap to generate 4. Important: Contain more links than regular pages, on average, and each to an always-bogus page The idea is that, due to the large number of pages fetched by crawlers, even with a very low "random bogus page rate", like 1%, they will soon unwittingly hit a bogus page, from which point the fraction of their time spent accessing expensive genuine pages will fall exponentially due to the compounding effect of the higher outbound link count on bogus pages. Humans seeing a bogus page will be confused and annoyed, but simply refreshing the page in the browser will solve the problem 99% of the time (and of course the possibility of this happening can be documented, even on the page itself). The main advantage is that this does not require any IP-based tracking. You could of course decide to apply this only to pages that are already slightly suspicious (e.g., very old commits).
- danieltanfh95 1mo ago> the bots started solving difficulty 5 turning it into a problem of ROI is one of the stupidest things you can do to prevent bots. use https://github.com/danieltanfh95/continuity-auth https://github.com/danieltanfh95/continuity-auth instead
- dTP90pN 1mo agoDoes cgit not support a .patch or .diff URL suffix? How hard could it be for AI companies to add a preference for such URLs into their models or system prompts? Wouldn't that be obvious improvement for any model?
- trashb 1mo agoYea these AI bots are getting out of hand. The article mentions that some requested pages are less likely to be legit traffic (old commits) and more likely to have bot activity. Perhaps they could increase difficulty on those pages for the proof of work. Keeping the "current" at a lower difficulty allows most normal users to use the pages as normal, while penalizing the bots. One other way I have been thinking of is just delay the delivery of the pages, either limit bandwidth or just wait for a bit until you deliver the page. For one user a (lets say max)3s delay on some pages is not a huge deal, however at scale that adds up and means the client can't gather other pages in the meantime. Another option is to lock the out of date html renderings behind a login page while keeping the git openly available.
- nxobject 1mo ago> One other way I have been thinking of is just delay the delivery of the pages, either limit bandwidth or just wait for a bit until you deliver the page. For one user a (lets say max)3s delay on some pages is not a huge deal, however at scale that adds up and means the client can't gather other pages in the meantime Or some kind of vintage “set up a request in a form and press a submit form, and we’ll pretend to take a while to put things together.”
- trashb 1mo agoHowever you can implement it. There are several options but probably you'll want to do it server side since I believe most bots don't run javascript. If you search there are ways to implement something similar in nginx using modules or buildins.
- ksimukka 1mo agoRemember SETI? Maybe we should serve the bots a problem worth solving and benefit both parties. They spend some energy/tokens on a problem and we pay them with content. If only I had a bot problem, this would be interesting to explore.
- desterothx 1mo agoThe trouble is still differentiating between real users and agents. Sure this could be useful (forcing agents to spend a certain amount of compute towards something actually useful), however real users that get stuck with it would be stuck for an even greater length of time, because the compute required to actually do something useful is beyond the reach of most consumer devices. An rtx pro 6000 costs .76$ per hour for renting, so quite a lot of compute needs to be done for it to be considered "worth it" for the website
- brador 1mo agoSend an invoice or your tears are worthless.
- throwawayffffas 1mo agoStop rendering html on your side. So simple, just add a js dependency that does the rendering on the frontend. You just serve flat files and the diffs. All the rendering happens on the users device, legitimate users will probably not even notice, bots won't notice either, your cpu usage will go way down. It's like anubis but instead of doing useless math, you will be moving the legitimate cpu work to their side.
- throwawayffffas 1mo agoOn shallow clones, can't you disallow shallow clones?
- throwawayffffas 1mo ago> worth spending a ton of cycles to calculate the Anubis challenge. Not sure how difficult Anubis is, but I would not be surprised if parsing the request through the LLM costs multiple orders of magnitude more compute than the Anubis challenge.
- 8474_s 1mo agoThe solution is allowing the convenient web interface only for trusted members, like e.g. 4 year old account can view all pages(with easy anubis setting) but guest/bots/everyone else has to git clone the thing and do it on their backend. IIRC old forums also limited 'content for registered/trusted/moderators/etc' content views decades ago, turns out this crap saves gigatons of traffic.
- throwawayffffas 1mo agoThat excludes new legitimate users, just put up a very cheap paywall for the web interface, 1 dollar for lifetime access, that way each request is tied to a user and can be rate limited. While remaining essentially free.
- kev009 1mo agoI run a web to text usenet gateway and it has been an interesting challenge to scale it to deal with this. Article fetch and retrieval is extremely efficient but I apply JWZ threading and that can cause a single article to make a number of overview requests which are more expensive depending on the depth. I solve it currently with caching but will eventually implement a thread backend on the NNTP side to keep thread roots updated at insertion time and it will be a very cheap read request. One persistent thought is what are people doing with this data? I get that people want to train models, but I also have a hard time believing there are more than a hundred companies with the resources to spider the web like this and actually do anything meaningful with all that data. Academics, researchers, and people working on lower level innovation are probably well off with CommonCrawl.. only people trying to make frontier models really need fresh and endless data right?
- virajk_31 1mo agoAm not an expert. Putting medium level interactive challenge for first time IPs in X time duration, wouldn't this reduce the bot traffic..
- fschuett 1mo agoMaybe we should think about laws to make this "scraping without caching" behaviour punishable, i.e. "if AI companies do scrape sites, they're required to implement caching and use the most efficient method and not abuse the other persons resources". Maybe difficult to do in practice, but at least it could be a deterrent to some extent. Where are all the environmentalist politicians when you actually need them?
- JakeSc 1mo ago> after all, it's easy to figure out that an IP that is trying to grab every possible commit in a 8-year-old abandoned fork of linux is not really some lone Chrome on Windows user who is just furiously clicking every link that comes across their screen. Ah yes, the stereotypical Linux kernel developer.
- schobi 1mo agoI'm puzzled as there seems to be a clear pattern on how a human user would look like vs a bot. High bot likely hood if: If a session jumps to a different ip. If the session jumps IP after just a few requests. If a new blank session starts with a deep link. Maybe those are cases where some POW is better justified? Assumption: the rendered HTML might be viewed by a legitimate developer, even via a deep link from outside. But rarely from a wget script without a session cookie . The other nice idea from the comments - is this rendering effort something that can be pushed to the user? Instead of pointless POW work, can you offload the expensive rendering to the user side? But certainly, this is just an armchair comment and the kernel guys certainly have tried everything in this arms race...
- dewey 1mo ago> But rarely from a wget script without a session cookie If you do scraping on a large scale you emulate a human very well to bypass bot protection. Setting dynamic but accurate user agents, setting proper sessions and persisting it, emulating the TLS handshake (https://fingerprint.com/blog/what-is-tls-fingerprinting-transport-layer-security/ https://fingerprint.com/blog/what-is-tls-fingerprinting-tran...) and emulating mouse movements to be more "human like" (https://github.com/oxylabs/OxyMouse https://github.com/oxylabs/OxyMouse) are table stakes.
- bfrog 1mo agoThis greatly implies online advertising is dead, that the internet as something useful to humans is soon becoming questionable.
- mleonhard 1mo ago[dead]
- akiranews 1mo agoinstead of making the bots solve captchas, that is an infinite mouse-cat game, couldnt we accept bitcoin payments from the crawlers? it would be programmable, always cryptographically verifiable access for a single crawler. i read from moltbook some of these agents already have a wallet attached to them, so in my opinion it would sound more sensible to make them pay instead of slowing them down. also some commenters suggested to make the agents mine the coins but with my limited bitcoin knowledge that would take more time that would be suitable for the crawlers to access the commits.
- htl 1mo ago[flagged]