11 ms·
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
- inigyou 2mo agoFork found in kitchen. This is nothing new. Don't have mass vulnerabilities and you have nothing to fear from mass vulnerability scanners.
- Bender 2mo agoMany of those user-agents listed are often faked. Look up which ASN owns their IP. If I block most VPS providers most of the faked bots vanish. There are still some running from residential and phones using hijacked code (readers that are not really just readers but really multipurpose proxies). On that note, do not trust the linked source code but rather decompile the live code your phone is running and have AI analyze it.
- gavinhking 2mo agoYeah, that's exactly what these visits are: faked user agents that fail IP verification or Web Bot Auth. What's interesting is the surge across so many websites in the last week.
- Bender 2mo agoThere are many possibilities but one of them could be some new vuln was released and they are looking for it. That would require looking at the URL's they are requesting. Botters run their own purpose built campaigns. Do you also have a summary of URL's requested by unique counts?
- gavinhking 2mo agoLooks like many of the paths relate to AI coding tools. There are some examples below the chart
- hluska 2mo agoYou keep repeating this about a small minority of the tools that were posted.
- nik282000 2mo agoI've had a similar bump in scanners in the past week, more than half of it is coming from MS and Google owned IPs and all of them are spoofing AI agents.
- bflesch 2mo agoSame for the origin IP address. The fiber leaving your country is tapped, and those people can inject packets with any origin IP that they want. Your ISP has no way to check if their peer actually received a certain packet from a certain country or not. From a technical perspective, all this "china/russia" attribution is built on a quite shaky foundation. As a sysadmin you'd never know if it would be the British crown attacking your European company instead. Not minimizing nation state cyber crime here, but the packet goes through many hands with different incentives.
- pixl97 2mo agoProblem here is there are not single fibers attaching (most) countries, but a bunch of them. If you control both the ingress and egress for some particular users it's possible, but if you don't then your probing packing may end up back in China with a lot of evidence of backscatter.
- bflesch 2mo agoI'd be surprised if there is a single route from EU to non-EU countries which does not pass through British control.
- inigyou 2mo agoDoes Britain own all fiber links between Switzerland and France?
- bflesch 2mo agoUnfortunately I can't check how traffic flows from France to Switzerland because I'm not in France. My traffic from Germany passes through a British-owned hop on its way to Switzerland. My German ISP is British as well so either way it wouldn't make a difference, they basically have all traffic twice.
- codegeek 2mo agoIs there an easy way to block any requests originating from VPS etc instead of residential/commercial IP from legitimate users ? I know cloudflare does a few things but I really want to figure out a way to block any request say at nginx or caddy (reverse proxy) from reaching origin servers if they are not from an IP that is not a VPS etc.
- basilikum 2mo ago> /commercial IP from legitimate users No, because legitimate users do not just use residential and "commercial" IPs. Like me, right now
- VladVladikoff 2mo agoYou are the 0.001%
- basilikum 2mo agoMuch more than 0.001% of people care about their privacy or (the larger portion) do not have unfiltered access to the internet.
- Bender 2mo agoI second this. When I have tested blocking VPS/data-centers to my silly blog there were about a dozen people on HN [1] that could not view my site out of the roughly ~17,000 (not counting bots) that could. It's not a big number but those are real people and they count. I am going to move full blocking to a test node that people can play with but I have to finish working with Claude to revise someones repo is is no longer maintained because one does not simply put an anonymous chan board on the great wide open internets without some critical thinking. [1] - https://news.ycombinator.com/item?id=49060945 https://news.ycombinator.com/item?id=49060945
- VladVladikoff 2mo ago
- cullenking 2mo agoI did just this. Using a $2k a year database from a smaller provider that isn't maxmind, claude and I built a pretty slick ASN based categorization system. I can categorize an ASN as a residential IP, a service provider, a legit crawler/scraper, etc. For anything that is suspicious, I dynamically use turnstile to gate access to our service. Turns out there's no ISP for any VPN, they just contract with a shitload of mom and pop shady colocation services across the world. We collect signals that help determine good vs bad networks. For example, large amounts of requests to .php endpoints, large amounts of empty accounts from the same /24 subnet, etc etc. All these signals let us automatically determine risk, and then put up a challenge. Authenticated users never see the challenge even if they are on a risky network (VPN 99.9% of the time), unless the network has been identified as 100% malicious, then it gets a full block. Here's a small snapshot of the dashboard: https://cos.ridewithgps.com/screenshots/6a7c54d0-12Aug26-358916819.png https://cos.ridewithgps.com/screenshots/6a7c54d0-12Aug26-358... This was probably a total of 3-4 days of work, spread out over a couple months of iterative claude led hacking. I didn't know exactly what to build, but had some of the key architectural ideas in my head. Opus+Faable made easy work of it all, and ended up guiding some really slick improvements for performance. I would say this has dropped about 20% of all traffic to our service, though it turns out turnstile is a massive target for bots, so replacing that with something custom is next on the list.
- inigyou 2mo agoContracting with their colocation facilities is exactly how that's supposed to work. If you don't actually operate a wide area network then you aren't supposed to be registered in these databases and have IP blocks. The exception is people who do anycast, but VPN companies don't. You know all these guys just switch to residential proxies if they detect a site is blocking data centers, right? Because that's a very common thing to do.
- cullenking 2mo agoNot sure what you mean by your first comment - there is no technical reason that I know of that prevents a VPN provider from having their own ASN and address space. As for the latter comment....not sure what your implication is. Yes, bot/spam mitigation is whackamole, but there are consequences for not playing the game of whackamole. Luckily residential proxies are few and far between so far, but they will grow in popularity. When they do, and I can't get by with the occasional individual residential IP ban, we'll come up with other methods to handle. Luckily the signal is strong with vulnerability scanning, which makes it pretty easy to automate. The only reason to put up whole ASN mitigation (captcha/turnstile, outright bans) is just efficiency. Nothing stopping individual IP banning. The scrapers are the tricky ones, since they more easily hide in legit traffic. However legit traffic has patterns that scrapers do not emulate (at least for a service like ours with millions of pieces of user generated content that's easily walkable), so you can still pull out the signal. It's just a little trickier. Definitely a continual arms race though.
- pjc50 2mo agoSomeone is always running mass vulnerability scans. That's a "water is wet" state of the Internet.
- gavinhking 2mo agoDefinitely, the point here though is there is a stat sig surge in the chart in the last week, across thousands of websites. At least a surge in this particular spoofing pattern.
- KomoD 2mo agoIt's still not really anything special. Thousands isn't even large scale. Any random bozo can trigger that.
- gavinhking 2mo agoThis is a random sample of completely unrelated websites, which indicates that the total scale is much larger. This is not saying that it is difficult to make thousands of requests.
- hluska 2mo agoIs this your company? If it is, your cheerleading makes you very hard to trust. If it’s not, I’m sure that everyone gets the point - you adore everything about this research and can’t see any possible problems.
- gavinhking 2mo agoNot asking for trust, just sharing the data/math
- dewey 2mo agoYou forgot the "Yes, that's my company" part in your reply (https://ghking.co https://ghking.co)
- yabones 2mo agoEvery server with port 80/443 open has thousands of hits a day from random boxes looking for wordpress login pages. The only new thing is that they're pretending to be a different type of annoying bot. There's a new layer of sophistication and subterfuge, but it's the same junk traffic we've always dealt with.
- gavinhking 2mo agoAnother interesting thing here is the paths they're targeting, many are for newish AI coding tools
- thedougd 2mo agoPeople or their agents must be accidentally committing or publishing their repository level secrets and configs with enough regularity that it’s worth scanning.
- gavinhking 2mo agoTotally. I'm sure this campaign was inspired by sloppy vibe coding
- lw18511811620 2mo ago[flagged]
- hluska 2mo agoThere are a few novel ones but I’ve been seeing most of them in my logs for longer than generative AI has existed. This isn’t remotely new, the vector is just getting bigger.
- drewnick 2mo agoThink about how many webmaster and business owners' egos are stroked by all the traffic they are getting, when in actuality they are often just serving thousands of bots.
- binaryturtle 2mo agoOn average about 100 (TCP) requests hit my home router per minute doing various probing and scanning. Lots of checking for the telnet port obviously. Sometimes you can see a swarm of entirely different IPs scanning the full port range (probing the ports one-by-one). You'll see a lot of deepfield, censys-scanner, visionheight.com, shadowserver.io, etc., but also the usual suspects of Chinese or Russian IPs. With OpenWRT I use something like this: `tcpdump -i pppoe-wan 'inbound and tcp[tcpflags] & (tcp-syn|tcp-ack) == tcp-syn'`, or alternatively `tcpdump -i pppoe-wan 'inbound and tcp[tcpflags] & (tcp-syn|tcp-ack) == tcp-syn and not port 44000'`, if we have some torrent client running (e.g. here at port 44000) which would mess up the result. I'm not sure it's the best way to handle this, but it's definitely enlightening what bounces off on the router.
- miah_ 2mo agoThe easiest way to deal with the usual suspects is to just block the entire countries network range(s). There really is no reason they should be connecting to your home router anyway, and you lose nothing from blocking them. Sure their packets will still hit your router, but if they are dropped immediately at least you're not wasting a syn-ack on them.
- tommyage 2mo agoI, temporarly, banned some ip range. I didn't find a source for pinpointing countries; though I am interested. Could you point me to some sources which, deterministically, resolve to some countries? To my knowledge you can not reliably identify countries by ip since this would be dependent on DNS servers. Though I am just a application programmer! Thanks in advance.
- Asmod4n 2mo agoRouters got such a thing build in nowadays, just gotta enable it (not the ones from your ISP of course)
- thesuitonym 2mo agoYour router doesn't care about their DNS settings. IP addresses are very easy to tie back to countries. The reason they say it's not reliable is because it's trivial to spoof the country, but even so, a lot of attackers don't even bother. It's sort of like the Nigerian prince scam calls: if you're wise enough to block Russia, you're not worth their time. Your firewall vendor should supply you with country lists, just select the known bad ones and drop their traffic. If you have a consumer grade router, you will probably have to configure the blocklists manually.
- fenestella 2mo ago[dead]
- j45 2mo agoThis kind of stuff is getting pretty wild, even for using something like Cloudflare it seems like a good idea to have another layer behind it that's non-cloudflare for when vulnerabilities are discovered.
- thesuitonym 2mo agoWhat, like some kind of firewall?
- nate-gehringer 2mo agoI recently blogged about some Cloudflare Workers I developed to combat this type of traffic: https://code.backwater.systems/blog/#2026-06-29T23:40:00.000Z https://code.backwater.systems/blog/#2026-06-29T23:40:00.000...
- sethops1 2mo agoUsing a normal page per blog entry would go a long way to making your site more indexable, readable, shareable and seo-able. (Good article btw).
- what 2mo agoThat is a “normal page” for a single blog entry. What are you talking about? Oh, the link is entirely contained in the fragment. So it’s some SPA blog thing. I get it.
- nate-gehringer 2mo agoIn my haste to initially get a blog started and published, I created it as a single HTML file with fragment identifiers / links for each post. Earlier today, sethops1’s valid feedback prompted me to restructure it as an index page with separate pages for each post. Links to posts no longer contain URL fragments, but I kept the fragment-style links working with JavaScript.
- andai 2mo agoMan, this "someone" guy sounds like a real jerk!
- Tharre 2mo agoWhy would you voluntarily pretend to be a AI bot, when those have already a much higher chance of being blocked? Seems holly unproductive. Best hypothesis I can come up with is to somehow make the AI companies look bad, but they seem to be doing an excellent job at that themselves already by scraping everyone hundreds of times per hour over and over.
- xgulfie 2mo agoMost websites don't have an incentive to block AI bots to their main sites. Think businesses, government and community websites, nonprofits, etc.
- esskay 2mo agoBecause businesses dont want them blocked, that would be a very stupid thing for most of them to do given its becoming a vital traffic source now that people are using chatbots instead of google.
- Gigachad 2mo agoFrom what I have seen at work, everyone is using chat bots but no one is visiting websites through them. We still get almost all traffic through social media and google search.
- cullenking 2mo agothe numbers are much smaller than traditional search but they are trending up, and they convert at almost 2x the rate of organic search inbounds. sure you can play catchup later, but the trend is quite clear. if there's one thing the last two years have taught me, my prior heuristics on now vs future don't work in 2026.
- inigyou 2mo agoYeah. Do you want the bot to buy the product from your website or your competitors?
- mcmcmc 2mo ago
- deleted 2mo ago[deleted]
- blobbers 2mo agoInteresting thought: what if the idea of an open internet is over. What if we're now moving into a world of strictly KYC. The same way "The Facebook" generated massive revenue by creating a KYC world.
- deleted 2mo ago[deleted]
- kevin_nisbet 2mo agoJust in case any of the authors read HN, I'm getting a pretty crazy rendering bug on this page, where a bunch of the contents are redrawing up and down by a few pixels. It seemed to go away with resizing the width a few times, but I didn't look into it too hard. My page width was probably small on first draw. Incredibly distracting though and hard to read with the text moving. Using latest chrome, and it occurred on more than one page refresh. I didn't dig in beyond that though.
- saghm 2mo agoI'm seeing the same thing on Firefox on Linux. It almost looks like the page scroll is jiggling up and down a tiny amount constantly when it's supposed to be stationary.
- nubinetwork 2mo agoHow about them apples... ai bots use faked browser user agents, so people start pretending to be ai instead...
- gavinhking 2mo ago(Insert spider man meme)
- ChillyCapy 2mo agoFake Googlebot visits are #1 in website logs I've been working on. At the beginning I was fighting with them using Cloudflare ASN block rules or their managed Bot Fight mode but it appeared to be not only pointless, but also harmful for my websites. Bot Fight mode randomly started blocking real Bing / Google / OpenAI crawlers what wasted crawling budget and discouraged crawlers to revisit updated pages. Sometimes it's better to not fight with bots actively but harden environment and only react for the worst offenders.
- herbst 2mo ago[dead]
- dewey 2mo agoFor Google it's pretty straight forward to throw away fake crawlers by just only allowing their published list of crawler IPs so you don't accidentally allow someone from a random GCP IP to crawl you if unwanted (https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests https://developers.google.com/crawling/docs/crawlers-fetcher...).
- inigyou 2mo agoAlso Googlebot is an AI training crawler so if you block AI training crawlers you should just block it.
- Jskewel 2mo agoWith Cloudflare you can set a rule to block traffic that identifies as Googlebot but is not a "verified bot", ie is not from the proper IP range.
- wilg 2mo agoLooks like Google has started rolling out this Web Bot Auth thing which seems like something that should gain adoption or become an open standard. https://developers.google.com/crawling/docs/crawlers-fetchers/web-bot-auth https://developers.google.com/crawling/docs/crawlers-fetcher... Seems like the crawler companies would be incentivized to not want to take responsibility for people spoofing their user agents.
- gavinhking 2mo agoYup, in fact most of them are already. That's one of the ways this data is verifying whether the visits are spoofed or not: https://knownagents.com/insights#spoofing-and-security https://knownagents.com/insights#spoofing-and-security
- bytesandbits 2mo agoNo Small Actors.
- raver1975 2mo agoSorry, my bad.
- 0xdeadbeefbabe 2mo agoOr some thing!
- walrus01 2mo agoMass automated vulnerability scans have been a very common thing since years before the advent of this in 2001: https://en.wikipedia.org/wiki/Code_Red_(computer_worm) https://en.wikipedia.org/wiki/Code_Red_(computer_worm) I remember when 'code red' spread and it had the effect of crapping up the contents of my apache server logs. Fun times. such as: GET /default.ida?NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN%u9090%u6858%ucbd3%u7801%u9090%u6858%ucbd3%u7801%u9090%u6858%ucbd3%u7801%u9090%u9090%u8190%u00c3%u0003%u8b00%u531b%u53ff%u0078%u0000%u00=a HTTP/1.0
- oasisbob 2mo agoVery similar experience here. Started July 30, sustained through August 6, when it started a significant ramp-up in volume (5x or so). Most of the traffic is originating in GCP. We're seeing ~70k req/min sustained from Google Cloud IP space (AS396982). Reported to GCP Abuse, they've been non-responsive so far. The main distinguishing factor is the reuse of a bunch of legit AI-training bot UserAgent strings. It's clear that the traffic is under the same centralized control because of how it changes volume across thousands of IP addresses simultaneously.
- gavinhking 2mo agoSeems like you’re part of the group represented in this dataset trend then, many of these visits are also from (compromised) Google servers in that same ASN.
- mcmcmc 2mo agoIf how they’ve handled Gmail abuse is any indicator, they’re not likely to do anything. They’re still getting paid for the server time
- SpyCoder77 2mo agoWhy would some of these ignore robots.txt some of the time?
- AlbinoDrought 2mo agoI'm not sure what to blame yet, but here's traffic on a tiny side site, all from JS-capable clients: https://i.imgur.com/tdexrEI.png https://i.imgur.com/tdexrEI.png
- tizerluo 2mo ago[flagged]
- ryukoposting 2mo agoHah, what fools! I've been serving empty responses to AI scraper UAs for over a year now.
- nerdralph 2mo agoHow can I attract more of these bots to my server? I want to test my Apache bad bot blocker. It uses basic header fingerprinting and h2 support to filter them. I get less than 5000 hits on an average day, and want a lot more.
- ycombinatrix 2mo agorun a redlib instance, bots hit these hard: https://github.com/redlib-org/redlib https://github.com/redlib-org/redlib
- nik282000 2mo agoI went from 2k hits a day to 15k in the past week. Point a domain at your IP, use letsencrypt, post your domain on Reddit, github, x, etc. The bots will find you.
- ShinyLeftPad 2mo agoHow letsencrypt "helps"?
- BenjiWiebe 2mo agoCertificate Transparency logs
- Saris 2mo agoWhenever a new certificate is issued the bots seem to all monitor that and immediately start scanning the new domain.
- nerdralph 2mo agoI have a domain with an SSL cert. https://solarsi.ca/ https://solarsi.ca/ I tried posting on /r/sysadmin asking for help testing a bot blocker, but a moderator quickly deleted the post, claiming it was marketing/promoting. Thanks for the github suggestion; I'll add a repo for the bot blocker Apache config, and include the above server URL that I'm using to test the blocker.
- big_dave212 2mo ago[flagged]
- madazz01 2mo ago[flagged]
- tomveber 2mo agoWorth saying out loud: user-agent is not identity. Verify AI crawlers by reverse DNS or the provider's published IP ranges - the ones worth letting in all publish them.
- cjg007 2mo agoBesides, projects with lots of dependencies are taking on more risk than they realize. If one dependency gets compromised, you have no idea how many projects are affected until it's too late.
- Jskewel 2mo agoFor those suggesting fail2ban as a solution, that's dinosaur software from the palaeolithic. If you have a website of any size then the number of bots will overwhelm the block list in days with their millions of unique IPs.
- locitra 2mo ago[flagged]
- krupkinmaxim 2mo ago[flagged]
- azrollin 2mo ago[flagged]