6 ms·
LWN is currently under the heaviest scraper attack seen yet
- zahlman 9mo agoIs it still ongoing? The thread appears to be over 24 hours old and as a quick test I had no issue loading the main page (which is as snappy and responsive as expected from a low-bandwidth site like LWN).
- jzb 9mo agoNot at the moment. It’s subsided for now.
- blibble 9mo agothe perverse incentive is if you ddos the website such that it shuts down, no other "AI" parasites can get the valuable data big tech incentivised to ddos... what a world they've built
- ronsor 9mo agoThis sounds like a conspiracy theory.
- MBCook 9mo agoI don’t think they’re saying that’s actually happening here, just that it could happen and is accidentally incentivized.
- pwdisswordfishy 9mo agoIf it's a conspiracy, it would be one where the Minimum Viable Conspirator Count is 1 (inclusive of one's own self). In that case, by that rubric literally anything that you conspire with yourself to accomplish (buying next week's groceries, making a turkey sandwich...) would also be a conspiracy.
- amlib 9mo agoThe dead internet theory also sounded unhinged and conspiracy theory-ish a decade or so ago... yet here we are.
- phkahler 9mo agoIts called pulling up the ladder behind you, or building a moat!
- NitpickLawyer 9mo agoUmm... what data? That's a very old newsletter-like site. Everything that's public on it has been long scraped and parsed by whoever needed it. There's 0 valuable data there for "parasites" to parasite off of. I also don't get the comments on the linked social site. IIUC the users posting there are somehow involved with kernel work, right? So they should know a thing or two about technical stuff? How / why are they so convinced that the big bad AI baddies are scraping them, and not some miss-configured thing that someone or another built? Is this their first time? Again, there's nothing there that hasn't been indexed dozens of times already. And... sorry to say it, but neither newsletters nor the 1-3 comments on each article are exactly "prime data" for any kind of training. These people have gone full tinfoil hat and spewing hate isn't doing them any favours.
- MBCook 9mo agoI don’t think they were talking about LWN specifically but just in general.
- homebrewer 9mo agoBecause it started in 2022 and hasn't subsided since? This is just the latest iteration of "AI" scrapers destroying the site, and the worst one yet. https://lwn.net/Articles/1008897 https://lwn.net/Articles/1008897 Your nonsense about LWN being a "newsletter" and having "zero valuable data" isn't doing you any favors. It is the prime source of information about Linux kernel development, and Linux development in general. "AI" cancer scraping the same thing over and over and over again is not news for anybody even with a cursory interest in this subject. They've been doing it for years.
- NitpickLawyer 9mo ago> LWN.net is a reader-supported news site I mean... Again, the site is so old that anything worth while is already in cc or any number of crawls. I am not saying they weren't scraped. I'm saying they likely weren't scraped by the bad AI people. And certainly not by AI companies trying to limit others from accessing that data (as the person who I replied to stated).
- gulugawa 9mo agoI've had luck blocking scrapers by overwriting JavaScript methods " a.getElementsByTagName = function (...args) {//Clear page content}" One can also hide components inside Shadow DOM to make it harder to scrape. However, these methods will interfere with automated testing tools such as Playwright and Selenium. Also, search engine indexing is likely to be affected.
- bogwog 9mo agoThis is a fun idea, especially if you make those functions procedurally generate garbage to get them stuck
- TurdF3rguson 9mo agoYou think you've had luck. The truth is you have no idea of knowing if this ever had any effect at all.
- chrisjj 9mo agoSo which is it? DDOS attack or "AI" scrapers?
- fabian2k 9mo agoSufficiently aggressive and inconsiderate scraping is indistinguishable from a DDOS attack.
- Y-bar 9mo agoA sufficiently stupid and egregious AI scraper is indistinguishable from a DDOS attack. Edit: Fabian2k was ten seconds ahead. Damn!
- TurdF3rguson 9mo agoScrapers because DDOS implies that it's malicious rather than accidental and there's no reason to think that.
- chrisjj 9mo agoRight, so probably the site should not be claiming "It is a DDOS attack".
- jacquesm 9mo agoAI allows companies to resell open source code as if they wrote it themselves doing an end run around all license terms. This is a major problem. Of course they're not going to stop at just code. They need all the rest of it as well.
- zipy124 9mo agoFrom the creators of easy money laundering (crypto bros), we now bring you easy money laundering 2: intellectual property laundering, coming to a theatre near you soon!
- gruez 9mo ago>From the creators of easy money laundering (crypto bros), Is there even any evidence that "crypto bros" and "AI bros" are even the same set of people other than being vaguely "tech" and hated by HN? At best you have someone like Altman who founded openai and had a crypto project (worldcoin), but the latter was approximately used by nobody. What about everyone else? Did Ilya Sutskever have a shitcoin a few years ago? Maybe Changpeng Zhao has an AI lab?
- themafia 9mo ago> and had a crypto project (worldcoin) That was a biometric surveillance project disguised as a crypto project. > Is there even any evidence that "crypto bros" and "AI bros" are even the same set of people No, the "AI" people are far worse. I always had a choice to /not/ use crypto. The "AI" people want to hamfistedly shove their flawed investment into every product under the sun.
- prussia 9mo agoI think the worst of the crypto people are probably the worst of the AI people too. Power/money-hungry grifters naturally move on to the most profitable grift when the old one peters out.
- pkaeding 9mo ago
- blakesterz 9mo ago"It is a DDOS attack involving tens of thousands of addresses" It is amazing just how distributed some of these things are. Even on the small sites that I help host we see these types of attacks from very large numbers of diverse IPs. I'd love to know how these are being run.
- smitty1e 9mo agoCall it a "Distributed Intelligence Logic Denial Of Service" (DILDOS) attack both to name it distinctly and characterize the source.
- random1234user 9mo agoMight as well call it "Artificial Intelligence Distributed Intelligence Logic Denial Of Service" (AIDILDOS) sounds about right.
- PaulDavisThe1st 9mo agoanother reference point: we've had well over 1M unique IP addresses hit git.ardour.org as part of stupid as hell git scraping effort. 1M !!!
- wongarsu 9mo agoThere are plenty of providers selling "residential proxies", distributing your crawler traffic through thousands of residential IPs. BrightData is probably the biggest, but its a big and growing market. And if you don't care about the "residential" part you can get proxies with data center IPs for much cheaper from the same providers. But those are easily blocked
- quectophoton 9mo agoAnd how do you get those residential IP addresses? Well, you just need people to install your browser extension. Or your proprietary web browser. Or your mobile app. Or your nice MCP. Maybe get them to add your PPA repository so they automatically install your sneakily-overriden package the next time they upgrade their system. Anything goes as long as your software has access to outgoing TCP port 443, which almost nobody blocks, so even if it's being run from within a Docker container or a VM it probably doesn't affect you.
- tedivm 9mo agoI solved this problem for my blog by simply not being interesting.
- fancyfredbot 9mo agoIf you can bore an LLM that's exciting.
- chuckadams 9mo agoBore-a-Bot, the new service from Confuse-a-Cat.
- antod 9mo agoThat sounds like Elon merging xAI and The Boring Company
- sandworm101 9mo agoI would rather setup a "shadow" site designed only for LLMs. I would stuff it with ao much insanity that Grok would not be able to leave. How about a billion blog post where every use of "American" is replaced with "Canadian". By the time im done, grok will be spouting conspiracy theories about the decline of the strategic bacon reserve.
- antod 9mo ago> By the time im done, grok will be spouting conspiracy theories about the decline of the strategic bacon reserve. grok will blame the zionists rather than the freemasons for that one.
- naiv 9mo agoTIL about Git Brag because of your blog. It is interesting.
- deleted 9mo ago[deleted]
- fancyfredbot 9mo agoWho are these agressive scrapers run by? It is difficult to figure out the incentives here. Why would anyone want to pull data from LWN (or any other site) at a rate which would cause a DDOS like attack? If I run a big data hungry AI lab consuming training data at 100Gb/s it's much much easier to scrape 10,000 sites at 10Mb/s than DDOS a smaller number of sites with more traffic. Of course the big labs want this data but why would they risk the reputational damage of overloading popular sites in order to pull it in an hour instead of a day or two?
- kylehotchkiss 9mo agochina (alibaba and tencent)
- fancyfredbot 9mo agoI'm not at all sure alibaba or tencent would actually want to DDOS LWN or any other popular website. They may face less reputational damage than say Google or OpenAI would but I expect LWN has Chinese readers who would look dimly on this sort of thing. Some of those readers probably work for Alibaba and Tencent. I'm not necessarily saying they wouldn't do it if there was some incentive to do so but I don't see the upside for them.
- philipkglass 9mo agoI don't think that most of them are from big-name companies. I run a personal web site that has been periodically overwhelmed by scrapers, prompting me to update my robots.txt with more disallows. The only big AI company I recognized by name was OpenAI's GPTBot. Most of them are from small companies that I'm only hearing of for the first time when I look at their user agents in the Apache logs. Probably the shadiest organizations aren't even identifying their requests with a unique user agent. As for why a lot of dumb bots are interested in my web pages now, when they're already available through Common Crawl, I don't know.
- iamnothere 9mo agoMaybe someone is putting out public “scraper lists” that small companies or even individuals can use to find potentially useful targets, perhaps with some common scraper tool they are using? That could explain it? I am also mystified by this.
- bloppe 9mo agoI'm curious how they concluded this was done to scrape for AI training. If the traffic was easily distinguishable from regular users, they would be able to firewall it. If it was not, then how can they be sure it wasn't just a regular old malicious DDOS? Happens way more often than you might think. Sometimes a poorly-managed botnet can even misfire.
- MBCook 9mo agoWhy would anyone ever DDOS them? They’ve been around for about three decades now, I don’t know if they’ve ever had a DDOS attack before the AI crawling started.
- iamnothere 9mo agoI am starting to think these are not just AI scrapers blindly seeking out data. All kinds of FOSS sites including low volume forums and blogs have been under this kind of persistent pressure for a while now. Given the cost involved in maintaining this kind of widespread constant scraping, the economics don’t seem to line up. Surely even big budget projects would adjust their scraping rates based on how many changes they see on a given site. At scale this could save a lot of money and would reduce the chance of blocking. I haven’t heard of the same attacks facing (for instance) niche hobby communities. Does anyone know if those sites are facing the same scale of attacks? Is there any chance that this is a deniable attack intended to disrupt the tech industry, or even the FOSS community in particular, with training data gathered as a side benefit? I’m just struggling to understand how the economics can work here.
- zomiaen 9mo agoHow many of these scrapers are written by AI by data-science folks who don't remotely care how often they're hitting the sites, and is data they wouldn't even think to give or ask the LLM about?
- iamnothere 9mo agoBut does that explain all of the various scrapers doing the same thing across the same set of sites? And again, the sheer bandwidth and CPU time involved should eventually bother the bean counters. I did think of a couple of possibilities: - Someone has a software package or list of sites out there that people are using instead of building their own scrapers, so everyone hits the same targets with the same pattern. - There are a bunch of companies chasing a (real or hoped for) “scraped data” market, perhaps overseas where overhead is lower, and there’s enough excess AI funding sloshing around that they able to scrape everything mindlessly for now. If this is the case then the problem should fix itself as funding gets tighter.
- TurdF3rguson 9mo agoMy theory on this one is some serial wantrepreneur came up with a business plan of scraping the archive and feeding it into a LLM to identify some vague opportunity. Then they paid some Fiverr / Upwork kid in India $200 to get the data. The good news is this website and any other can mitigate these things by moving to Cloudflare and it's free.
- 2OEH8eoCRo0 9mo agoWhen are we going to start suing these assholes? Why isn't anybody leveraging the legal system? You're all searching for technical solutions to a legal problem and fighting with one hand behind your back.
- seb1204 9mo agoIs it possible to attribute the attack to a company?
- 2OEH8eoCRo0 9mo agoNope. It's impossible to trace anything that happens on the internet.
- Havoc 9mo agoThat makes no sense. There is no reason for AI scrappers to use tens of thousands of IPs to scrape one site over and over. That just sounds like a classic DDOS.
- TurdF3rguson 9mo agoSure there is, scrapers do that to defeat throttling. 10,000 is less than 3 hours of scraping at 1 request per second.
- Havoc 9mo agoIt's not 10k requests, it's 10k IPs Having lots of IPs is helpful for scraping, but you don't need 10k. That's a botnet
- TurdF3rguson 9mo agoThe way it works is this: You can sign up for a proxy rotator service that works like a regular proxy except every request you make goes through a different ip address. Is that a botnet? Yes. Is it also typically used in a scraping project? Yes.
- Havoc 9mo agoYeah I know, I've done scrapping too. It can absolutely be that, but that requires a confluence of multiple factors - misconfigured scrapper hitting the site over and over, a big bot net like proxy setup that is way overkilled for scrapping, a setup sophisticated enough to do all that yet simultaneously stupid enough to not cope with a site is mostly text and a couple gigs at most and all that over extended timeframe without anyone realising their scrapper is stuck. Or alternative explanation: It's a DDOS
- TurdF3rguson 9mo agoExcept that I think it's clear that the motive was getting the data not taking the site offline. The evidence for that is that it stopped on its own without them doing anything to mitigate it. Also I don't know why you think this is sophisticated, it's probably 40 lines of Python code max.
- sgc 9mo agoCan somebody tell me what is a normal "cost of doing business" level of bot traffic these days? I have way too much bot traffic like everybody else, but I don't know if I am an outlier or just run of the mill. I get about 100k bot hits a day, presumably because I have about 350k pages on my site.
- kay_o 9mo agoEsports vertical: I get about 5-20b bot hits per day (unwanted; includes both IA, brute forcer, "security" scanners, wp-admin/ requests), 1.5m google spider (search; respectful of crawl delay), and about 50-100m human (largely mobile). For unwanted bots I serve incorrect information -- it's online gaming match history without much text so requests flagged as unwanted bots will, instead of heavy database queries, get plausibly random numbers -- seeded by the user so they stay stable -- KDA, win/loss rates, rankings. A few dozen million distinct pages but they are numeric stats for user profiles, match stats with little to none paragraph form of text.
- sgc 9mo agoWell you must be an outlier! 100-200x bot to human traffic is a lot. AI bots likely focus on stats / technical info if they happen to be tuned to discern, and there is money in sports stats; so you are a target.
- samtrack2019 9mo agoI gave up and moved to cloudfare for my static blog, it was too much of a pain to track the spammers. (i am not proud of using big$corp)
- xacky 9mo agoResidential proxies need to be classed as malware and added to antivirus definitions and kicked off of app stores. Mass ban evasion needs to be cracked down on.