26 ms·
Ban me at the IP level if you don't like me
- Slava_Propanei 1y ago[dead]
- _def 1y agoI've seen blocks like that for e.g. alibaba cloud. It's sad indeed, but it can be really difficult to handle aggressive scrapers.
- Etheryte 1y agoOne starts to wonder, at what point might it be actually feasible to do it the other way around, by whitelisting IP ranges. I could see this happening as a community effort, similar to adblocker list curation etc.
- worthless-trash 1y agoI admin a few local business sites.. I whitelist all the countries isps and the strangeness in the logs and attack counts have gone down. Google indexes in country, as does a few other search engines.. Would recommend.
- coffee_am 1y agoIs there a public curated list of "good ips" to whitelist ?
- worthless-trash 1y agoSo, its relatively easy because there is limited ISP's in my country. I imagine its a much harder option for the US. I looked at all the IP ranges delegated by APNIC, along with every local ISP that I could find, unioned this with https://lite.ip2location.com/australia-ip-address-ranges https://lite.ip2location.com/australia-ip-address-ranges And so far i've not had any complaints. and I think that I have most of them. At some time in the future, i'll start including https://github.com/ebrasha/cidr-ip-ranges-by-country https://github.com/ebrasha/cidr-ip-ranges-by-country
- partyguy 1y ago> Is there a public curated list of "good ips" to whitelist ? https://github.com/AnTheMaker/GoodBots https://github.com/AnTheMaker/GoodBots
- plaguna 1y ago[flagged]
- valvix 1y ago[flagged]
- plaguna 1y ago[flagged]
- ThrowMeAway1618 1y ago>There is no need to disagree on such strongly worded statements. What's the bigoted history of those terms? from here[0]: "The English dramatist Philip Massinger used the phrase "black list" in his 1639 tragedy The Unnatural Combat.[2] "After the restoration of the English monarchy brought Charles II of England to the throne in 1660, a list of regicides named those to be punished for the execution of his father.[3] The state papers of Charles II say "If any innocent soul be found in this black list, let him not be offended at me, but consider whether some mistaken principle or interest may not have misled him to vote".[4] In a 1676 history of the events leading up to the Restoration, James Heath (a supporter of Charles II) alleged that Parliament had passed an Act requiring the sale of estates, "And into this black list the Earl of Derby was now put, and other unfortunate Royalists".[5]" Are you an enemy of Charles II? Is that what the problem is? [0] https://en.wikipedia.org/wiki/Blacklisting#Origins_of_the_term https://en.wikipedia.org/wiki/Blacklisting#Origins_of_the_te...
- defrost 1y ago[flagged]
- ThrowMeAway1618 1y agoThe origin of the term 'black list' had absolutely nothing to do with the melanin content of anyone. In fact, when that term was coined, it had nothing to do with the melanin content of anyone. It was a list of the enemies of Charles II. That's why I posted that. I'd also point out that in my lifetime, folks with darker skin called themselves black and proudly so. As Mr. Brown[0][1] will unambiguously tell you. Regardless, claiming that a term for the property of absorbing visible light is bigoted, to every use of such a term is ridiculous on its face. By your logic, if I wear black socks, I'm a bigot? Or am only a bigot if I actually refer to those socks as "black." Should I use "socks of color" so as not to be a bigot? If I like that little black dress, I'm a bigot as well? Or only if I say "I like that little black dress?" Look. I get it. Melanin content is worthless as a determinant of the value of a human. And anyone who thinks otherwise is sorely and sadly mistaken. It's important to let folks know that there's only one race of sentient primates on this planet -- Homo Sapiens. What's more, we are all, no matter where we come from, incredibly closely related from a genetic standpoint. The history of bigotry, murder and enslavement by and to our fellow humans is long, brutal and disgusting. But nitpicking terms (like black list) that never had anything to do with that bigotry seems performative at best. As I mentioned above, do you also make such complaints about black socks or shoes? Black dresses? Black foregrounds/backgrounds? If not, why not? That's not a rhetorical question. [0] https://www.youtube.com/watch?v=oM1_tJ6a2Kw https://www.youtube.com/watch?v=oM1_tJ6a2Kw [1] https://www.azlyrics.com/lyrics/jamesbrown/sayitloudimblackandimproud.html https://www.azlyrics.com/lyrics/jamesbrown/sayitloudimblacka...
- ygritte 1y agoCame here to say something similar. The sheer amount of IP addresses one has to block to keep malware and bots at bay is becoming unmanageable.
- immibis 1y agoCan you explain more about blocking malware as opposed to bots?
- ygritte 1y agoNo opposition. Just block the IP address.
- leviathant 1y agoKnowing my audience, I've blocked entire countries to stop the pain. Even that was a bit of whack-a-mole. Blocking China cooled off the traffic for a few days, then it came roaring back via Singapore. Blocked Singapore, had a reprieve for a while, and then it was India, with a vengeance. Cloudflare has been a godsend for protecting my crusty old forum from this malicious, wasteful behavior.
- bobbiechen 1y agoUnfortunately, well-behaved bots often have more stable IPs, while bad actors are happy to use residential proxies. If you ban a residential proxy IP you're likely to impact real users while the bad actor simply switches. Personally I don't think IP level network information will ever be effective without combining with other factors. Source: stopping attacks that involve thousands of IPs at my work.
- throwawayffffas 1y ago> If you ban a residential proxy IP you're likely to impact real users while the bad actor simply switches. Are you really? How likely do you think is a legit customer/user to be on the same IP as a residential proxy? Sure residential IPS get reused, but you can handle that by making the block last 6-8 hours, or a day or two.
- micahdeath 1y agoWe blocked AT&T Mobile once... You get lots of complaints that way and we only blocked them for an hour.
- richardwhiuk 1y agoIn these days of CGNAT, a residential IP is shared by multiple customers.
- immibis 1y agoVery likely. You can voluntarily run one to make ~$10/month in cryptocurrency. Many others are botnets. They aren't signing up for new internet connections solely to run proxies on.
- BLKNSLVR 1y agoBlocking a residential proxy doesn't sound like a bad idea to me. My single-layer thought process: If they're knowingly running a residential proxy then they'll likely know "the cost of doing business". If they're unknowingly running a residential proxy then blocking them might be a good way for them to find out they're unknowingly running a residential proxy and get their systems deloused.
- delusional 1y agoAt that point it almost sounds like we're doing "peering" agreements at the IP level. Would it make sense to have a class of ISPs that didn't peer with these "bad" network participants?
- JimDabell 1y agoIf this didn’t happen for spam, it’s not going to happen for crawlers.
- shortrounddev2 1y agoWhy not just ban all IP blocks assigned to cloud providers? Won't halt botnets but the IP range owned by AWS, GCP, etc is well known
- hnlmorg 1y agoBecause crawlers would then just use a different IP which isn’t owned by cloud vendors.
- jjayj 1y agoBut my work's VPN is in AWS, and HN and Reddit are sometimes helpful... Not sure what my point is here tbh. The internet sucks and I don't have a solution
- aorth 1y agoTricky to get a list of all cloud providers, all their networks, and then there are cases like CATO Networks Ltd and ZScaler, which are apparently enterprise security products that route clients traffic through their clouds "for security".
- lxgr 1y agoMany US companies do it already. It should be illegal, at least for companies that still charge me while I’m abroad and don’t offer me any other way of canceling service or getting support.
- withinboredom 1y agoI'm pretty sure I still owe t-mobile money. When I moved to the EU, we kept our old phone plans for awhile. Then, for whatever reason, the USD didn't make it to the USD account in time and we missed a payment. Then t-mobile cut off the service and you need to receive a text message to login to the account. Obviously, that wasn't possible. So, we lost the ability to even pay, even while using a VPN. We just decided to let it die, but I'm sure in t-mobile's eyes, I still owe them.
- thenthenthen 1y agoThis! Dealing with European services from China is also terrible. As is the other way around. Welcome to the intranet!
- thenthenthen 1y agoIn addition, my tencent and alicloud instances are also hammered to death by their own bots. Just to add a bit of perspective.
- partyguy 1y agoThat's what I'm trying to do here, PRs welcome: https://github.com/AnTheMaker/GoodBots https://github.com/AnTheMaker/GoodBots
- aorth 1y agoNoble effort. I might make some pull requests, though I kinda feel it's futile. I have my own list of "known good" networks.
- friendzis 1y agoIt's never either/or: you don't have to choose between white and black lists exclusively and most of the traffic is going to come from grey areas anyway. Say you whitelist an address/range and some systems detect "bad things". Now what? You remove that address/range from whitelist? Doo you distribute the removal to your peers? Do you communicate removal to the owner of unwhitelisted address/range? How does owner communicate dealing with the issue back? What if the owner of the range is hosting provider where they don't proactively control the content hosted, yet have robust anti-abuse mechanisms in place? And so on. Whitelist-only is a huge can of worms and whitelists works best with trusted partner you can maintain out-of-band communication with. Similarly blacklists work best with trusted partners, however to determine addresses/ranges that are more trouble than they are worth. And somewhere in the middle are grey zone addresses, e.g. ranges assigned to ISPs with CGNATs: you just cannot reliably label an individual address or even a range of addresses as strictly troublesome or strictly trustworthy by default. Implement blacklists on known bad actors, e.g. the whole of China and Russia, maybe even cloud providers. Implement whitelists for ranges you explicitly trust to have robust anti-abuse mechanisms, e.g. corporations with strictly internal hosts.
- jampa 1y agoThe Pokémon Go company tried that shortly after launch to block scraping. I remember they had three categories of IPs: - Blacklisted IP (Google Cloud, AWS, etc), those were always blocked - Untrusted IPs (residential IPs) were given some leeway, but quickly got to 429 if they started querying too much - Whitelisted IPs (IPV4 addresses are used legitimately by many people), for example, my current data plan tells me my IP is from 5 states over, so anything behind a CGNAT. You can probably guess what happens next. Most scrapers were thrown out, but the largest ones just got a modem device farm and ate the cost. They successfully prevented most users from scraping locally, but were quickly beaten by companies profiting from scraping. I think this was one of many bad decisions Pokémon Go made. Some casual players dropped because they didn't want to play without a map, while the hardcore players started paying for scraping, which hammered their servers even more.
- aorth 1y agoI have an ad hoc system that is similar, comprised of three lists of networks: known good, known bad, and data center networks. These are rate limited using a geo map in nginx for various expensive routes in my application. The known good list is IPs and ranges I know are good. The known bad list is specific bad actors. The data center networks list is updated periodically based on a list of ASNs belonging to data centers. There are a lot of problems with using ASNs, even for well-known data center operators. First, they update so often. Second, they often include massive subnets like /13(!), which can apparently overlap with routes announced by other networks, causing false positives. Third, I had been merging networks (to avoid overlaps causing problems in nginx) with something like https://github.com/projectdiscovery/mapcidr https://github.com/projectdiscovery/mapcidr but found that it also caused larger overlaps that introduced false positives from adjacent networks where apparently some legitimate users are. Lastly, I had seen suspicious traffic from data center operators like CATO Networks Ltd and ZScaler that are some kind of enterprise security products that route clients through their clouds. Blocking those resulted in some angry users in places I didn't expect... And none of the accounts for the residential ISPs that bots use to appear like legitimate users https://www.trendmicro.com/vinfo/us/security/news/vulnerabilities-and-exploits/a-closer-exploration-of-residential-proxies-and-captcha-breaking-services https://www.trendmicro.com/vinfo/us/security/news/vulnerabil....
- guesswho_ 1y ago[dead]
- lwansbrough 1y agoWe solved a lot of our problems by blocking all Chinese ASNs. Admittedly, not the friendliest solution, but there were so many issues originating from Chinese clients that it was easier to just ban the entire country. It's not like we can capitalize on commerce in China anyway, so I think it's a fairly pragmatic approach.
- lxgr 1y agoWhy stop there? Just block all non-US IPs! If it works for my health insurance company, essentially all streaming services (including not even being able to cancel service from abroad), and many banks, it’ll work for you as well. Surely bad actors wouldn’t use VPNs or botnets, and your customers never travel abroad?
- lwansbrough 1y agoDon't care, works fine for us.
- yupyupyups 1y agoAnd that's perfectly fine. Nothing is completely bulletproof anyway. If you manage to get rid of 90% of the problem then that's a good thing.
- lxgr 1y agoAnd if your competitor manages to do so without annoying the part of their customer base that occasionally leaves the country, everybody wins!
- yupyupyups 1y agoFair point, that's something to consider.
- ruszki 1y agoOkay, but this causes me about 90% of my major annoyances. Seriously. It’s almost always these stupid country restrictions. I was in UK. I wanted to buy a movie ticket there. Fuck me, because I have an Austrian ip address, because modern mobile backends pass your traffic through your home mobile operator. So I tried to use a VPN. Fuck me, VPN endpoints are blocked also. I wanted to buy a Belgian train ticket still from home. Cloudflare fuck me, because I’m too suspicious as a foreigner. It broke their whole API access, which was used by their site. I wanted to order something while I was in America at my friend’s place. Fuck me of course. Not just my IP was problematic, but my phone number too. And of course my bank card… and I just wanted to order a pizza. The most annoying is when your fucking app is restricted to your stupid country, and I should use it because your app is a public transport app. Lovely. And of course, there was that time when I moved to an other country… pointless country restrictions everywhere… they really helped. I remember the times when the saying was that the checkout process should be as frictionless as possible. That sentiment is long gone.
- deleted 1y ago[deleted]
- ta8645 1y agoIf ipv6 ever becomes a thing, it'll make blocking all that much harder.
- snerbles 1y agoFor ipv6 you just start nuking /64s and /48s if they're really rowdy.
- rnhmjoj 1y agoNo, it's really the same thing with just different (and more structured) prefix lengths. In IPv4 you usually block a single /32 address first, then a /24 block, etc. In IPv6 you start with a single /128 address, a single LAN is /64, an entire site is usually /56 (residential) or /48 (company), etc.
- withinboredom 1y agoHmmm... that isn't my experience: /128: single application /64: single computer /56: entire building /48: entire (digital) neighborhood
- rnhmjoj 1y agoA /64 is the smallest network on which you can run SLAAC, so almost all VLANs should use this. /56 and /48 for end users is what RIRs are recommending, in reality the prefixes are longer, because ISPs and hosting providers wants you to pay like IPv6 space is some scarse resource. [1]: https://www.ripe.net/publications/docs/ripe-690/ https://www.ripe.net/publications/docs/ripe-690/
- withinboredom 1y agoEveryone at my isp is issued a /56 (and as far as I can tell, the entire country is this way).
- Arnavion 1y agoNote that for the sake of blocking internet clients, there's no point blocking a /128. Just start at /64. Blocking a /128 is basically useless because of SLAAC.
- firefoxd 1y agoSince I posted an article here about using zip bombs [0], I'm flooded with bots. I'm constantly monitoring and tweaking my abuse detector, but this particular bot mentioned in the article seemed to be pointing to an RSS reader. I white listed it at first. But now that I gave it a second look, it's one of the most rampant bot on my blog. [0]: https://news.ycombinator.com/item?id=43826798 https://news.ycombinator.com/item?id=43826798
- dmurray 1y agoIf I had a shady web crawling bot and I implemented a feature for it to avoid zip bombs, I would probably also test it by aggressively crawling a site that is known to protect itself with hand-made zip bombs.
- sim7c00 1y agoalso protect yourself fromnsucking up fake generated content. i know some folks here like to feed them all sorts of 'data' . fun stuff :D
- Moru 1y agoRule number one: You do not talk about fight club.
- popcorncowboy 1y agoDark forest theory taking root.
- JdeBP 1y agoOne of the few manual deny-list entries that I have made was not for a Chinese company, but for the ASes of the U.S.A. subsidiary of a Chinese company. It just kept coming back again and again, quite rapidly, for a particular page that was 404. Not for any other related pages, mind. Not for the favicon, robots.txt, or even the enclosing pseudo-directory. Just that 1 page. Over and over. The directory structure had changed, and the page is now 1 level lower in the tree, correctly hyperlinked long since, in various sitemaps long since, and long since discovered by genuine HTTP clients. The URL? It now only exists in 1 place on the WWW according to Google. It was posted to Hacker News back in 2017. (My educated guess is that I am suffering from the page-preloading fallout from repeated robotic scraping of old Hacker News stuff by said U.S.A. subsidiary.)
- PeterStuer 1y agoFAFO from both sides. Not defending this bot at all. That said, the shenanigans some rogue or clueless webmasters are up to blocking legitimate and non intrusive or load causing M2M trafic is driving some projects into the arms of 'scrape services' that use far less considerate nor ethical means to get to the data you pay them for. IP blocking is useless if your sources are hundreds of thousands of people worldwide just playing a "free" game on their phone that once in a while on wifi fetches some webpages in the background for the game publisher's scraping as a service side revenue deal.
- ahtihn 1y agoWhat? Are you trying to say it's legitimate to want to scrape websites that are actively blocking you because you think you are "not intrusive"? And that this justifies paying for bad actors to do it for you? I can't believe the entitlement.
- PeterStuer 1y agoNo. I'm talking about literally legitimate, information that has to be public by law and/or regulation (typically gov stuff), in formats specifically meant for m2m consuption, and still blocked by clueless or malicious outsourced lowest bidder site managers. And no, I do not use those paid services, even though it would make it much easier.
- geocar 1y agoExactly. If someone can harm your website on accident, they can absolutely harm it on purpose. If you feel like you need to do anything at all, I would suggest treating it like any other denial-of-service vulnerability: Fix your server or your application. I can handle 100k clients on a single box, which equates to north of 8 billion daily impressions, and so I am happy to ignore bots and identify them offline in a way that doesn't reveal my methodologies any further than I absolutely have to.
- BLKNSLVR 1y ago> IP blocking is useless if your sources are hundreds of thousands of people worldwide just playing a "free" game on their phone that once in a while on wifi fetches some webpages in the background for the game publisher's scraping as a service side revenue deal. That's traffic I want to block, and that's behaviour that I want to punish / discourage. If a set of users get caught up in that, even when they've just been given recycled IP addresses, then there's more chance to bring the shitty 'scraping as a service' behaviour to light, thus to hopefully disinfect it. (opinion coming from someone definitely NOT hosting public information that must be accessible by the common populace - that's an issue requiring more nuance, but luckily has public funding behind it to develop nuanced solutions - and can just block China and Russia if it's serving a common populace outside of China and Russia).
- phplovesong 1y agoWe block China and Russia. DDOS attacks and other hack attempts went down by 95%. We have no chinese users/customers so in theory this does not effect business at all. Also russia is sanctioned and our russian userbase does not actually live in russia, so blocking russia did not effect users at all.
- mavamaarten 1y agoSame here. It sucks. But it's just cost vs reward at some point.
- praptak 1y agoHow did you choose where to get the IP addresses to block? I guess I'm mostly asking where this problem (i.e. "get all IPs for country X") is on the scale from "obviously solved" to "hard and you need to play catch up constantly". I did a quick search and found a few databases but none of them looks like the obvious winner.
- bakugo 1y agoMaxmind's GeoIP database is the industry standard, I believe. You can download a free version of it. If your site is behind cloudflare, blocking/challenging by country is a built-in feature.
- tietjens 1y agoThe common cloud platforms allow you to do geo-blocking.
- spc476 1y agoI used CYMRU <https://www.team-cymru.com/ip-asn-mapping https://www.team-cymru.com/ip-asn-mapping> to do the mapping for the post.
- preinheimer 1y agoMaxMind is very common, IPInfo is also good. https://ipinfo.io/developers/database-download https://ipinfo.io/developers/database-download If you want to test your IP blocks, we have servers on both China and Russia, we can try to take a screenshot from there to see what we get (free, no signup) https://testlocal.ly/ https://testlocal.ly/
- herbst 1y agoMore than half of my traffic is Bing, Claude and for whatever reason the Facebook bots. None of these are main main traffic drivers, just the main resource hogs. And the main reason when my site turns slow (usually an AI, microsoft or Facebook ignoring any common sense) China and co is only a very small portion of my malicious traffic. Gladly. It's usually US companies who disrespect my robots.txt and DNS rate limits who make me the most problems.
- devoutsalsa 1y agoThere are a lot of dumb questions, and I pose all of them to Claude. There's no infrastructure in place for this, but I would support some business model where LLM-of-choice compensates website operators for resources consumed by my super dumb questions. Like how content creators get paid when I watch with a YouTube Premium subscription. I doubt this is practical in practice.
- herbst 1y agoFor me it looks more like out of the control bots than average requests. For example a few days ago I blocked a few bots. Google was about 600 requests in 24 hours. Bing 1500, Facebook is mostly blocked right now, Claude with 3 different bot types was about 100k requests in the same time. There is no reason to query all my sub-sites, it's like a search engine with way to many theoretical pages. Facebook also did aggressively, daily indexing of way to many pages, using large IP ranges until I blocked it. I get like one user per week from them, no idea what they want. And bing, I learned, "simply" needs hard enforced rate limits it kinda learns to agree on.
- roguebloodrage 1y agoThis is everything I have for AS132203 (Tencent). It has your addresses plus others I have found and confirmed using ipinfo.io 43.131.0.0/18 43.129.32.0/20 101.32.0.0/20 101.32.102.0/23 101.32.104.0/21 101.32.112.0/23 101.32.112.0/24 101.32.114.0/23 101.32.116.0/23 101.32.118.0/23 101.32.120.0/23 101.32.122.0/23 101.32.124.0/23 101.32.126.0/23 101.32.128.0/23 101.32.130.0/23 101.32.13.0/24 101.32.132.0/22 101.32.132.0/24 101.32.136.0/21 101.32.140.0/24 101.32.144.0/20 101.32.160.0/20 101.32.16.0/20 101.32.17.0/24 101.32.176.0/20 101.32.192.0/20 101.32.208.0/20 101.32.224.0/22 101.32.228.0/22 101.32.232.0/22 101.32.236.0/23 101.32.238.0/23 101.32.240.0/20 101.32.32.0/20 101.32.48.0/20 101.32.64.0/20 101.32.78.0/23 101.32.80.0/20 101.32.84.0/24 101.32.85.0/24 101.32.86.0/24 101.32.87.0/24 101.32.88.0/24 101.32.89.0/24 101.32.90.0/24 101.32.91.0/24 101.32.94.0/23 101.32.96.0/20 101.33.0.0/23 101.33.100.0/22 101.33.10.0/23 101.33.10.0/24 101.33.104.0/21 101.33.11.0/24 101.33.112.0/22 101.33.116.0/22 101.33.120.0/21 101.33.128.0/22 101.33.132.0/22 101.33.136.0/22 101.33.140.0/22 101.33.14.0/24 101.33.144.0/22 101.33.148.0/22 101.33.15.0/24 101.33.152.0/22 101.33.156.0/22 101.33.160.0/22 101.33.164.0/22 101.33.168.0/22 101.33.17.0/24 101.33.172.0/22 101.33.176.0/22 101.33.180.0/22 101.33.18.0/23 101.33.184.0/22 101.33.188.0/22 101.33.24.0/24 101.33.25.0/24 101.33.26.0/23 101.33.30.0/23 101.33.32.0/21 101.33.40.0/24 101.33.4.0/23 101.33.41.0/24 101.33.42.0/23 101.33.44.0/22 101.33.48.0/22 101.33.52.0/22 101.33.56.0/22 101.33.60.0/22 101.33.64.0/19 101.33.64.0/23 101.33.96.0/22 103.52.216.0/22 103.52.216.0/23 103.52.218.0/23 103.7.28.0/24 103.7.29.0/24 103.7.30.0/24 103.7.31.0/24 43.130.0.0/18 43.130.64.0/18 43.130.128.0/19 43.130.160.0/19 43.132.192.0/18 43.133.64.0/19 43.134.128.0/18 43.135.0.0/18 43.135.64.0/18 43.135.192.0/19 43.153.0.0/18 43.153.192.0/18 43.154.64.0/18 43.154.128.0/18 43.154.192.0/18 43.155.0.0/18 43.155.128.0/18 43.156.192.0/18 43.157.0.0/18 43.157.64.0/18 43.157.128.0/18 43.159.128.0/19 43.163.64.0/18 43.164.192.0/18 43.165.128.0/18 43.166.128.0/18 43.166.224.0/19 49.51.132.0/23 49.51.140.0/23 49.51.166.0/23 119.28.64.0/19 119.28.128.0/20 129.226.160.0/19 150.109.32.0/19 150.109.96.0/19 170.106.32.0/19 170.106.176.0/20
- bigiain 1y agoFor anyone wondering how to do this (like me from a month or two back). Here's a useful tool/site: https://bgp.tools https://bgp.tools You can feed it an ip address to get an AS ("Autonomous System"), then ask it for all prefixes associated with that AS. I fed it that first ip address from that list (43.131.0.0) and it showed my the same Tencent owned AS132203, and it gives back all the prefixes they have here: https://bgp.tools/as/132203#prefixes https://bgp.tools/as/132203#prefixes (Looks like roguebloodrage might have missed at least the 1.12.x.x and 1.201.x.x prefixes?) I started searching about how to do that after reading a RachelByTheBay post where she wrote: Enough bad behavior from a host -> filter the host. Enough bad hosts in a netblock -> filter the netblock. Enough bad netblocks in an AS -> filter the AS. Think of it as an "AS death penalty", if you like. (from the last part of https://rachelbythebay.com/w/2025/06/29/feedback/ https://rachelbythebay.com/w/2025/06/29/feedback/ )
- sneak 1y agoI feel like people seem to forget that an HTTP request is, after all, a request. When you serve a webpage to a client, you are consenting to that interaction with a voluntary response. You can blunt instrument 403 geoblock entire countries if you want, or any user agent, or any netblock or ASN. It’s entirely up to you and it’s your own server and nobody will be legitimately mad at you. You can rate limit IPs to x responses per day or per hour or per week, whatever you like. This whole AI scraper panic is so incredibly overblown. I’m currently working on a sniffer that tracks all inbound TCP connections and UDP/ICMP traffic and can trigger firewall rule addition/removal based on traffic attributes (such as firewalling or rate limiting all traffic from certain ASNs or countries) without actually having to be a reverse proxy in the HTTP flow. That way your in-kernel tables don’t need to be huge and they can just dynamically be adjusted from userspace in response to actual observed traffic.
- worthless-trash 1y ago> This whole AI scraper panic is so incredibly overblown. The problem is that its eating into peoples costs, and if you're not concerned with money, I'm just asking, can you send me $50.00 USD ?
- sneak 1y agoIf people don’t want to spend the money serving the requests, then their own servers are misconfigured because responding is optional.
- worthless-trash 1y agoSo, that is a no on the fifty? When AI can now register and break captures on your site to login, how do I compete with this arms race of defeating my protection from AI ?
- Grimblewald 1y agoIt sure is, now the problems i'd like to respond to legitimate users, I dont care for bots, if theres idle resources why not. Care to share how I can make that happen given scrapers are hellbent on ignoring any rules / agreements on how to conduct themselves?
- znpy 1y agoOh i recognise those ip addresses… they gave us quite an headache a while ago
- yumraj 1y agoWouldn't it be better, if there's an easy way, to just feed such bots shit data instead of blocking them. I know it's easier to block and saves compute and bandwidth, but perhaps feeding them shit data at scale would be a much better longer term solution.
- throwawayffffas 1y agoNo serving shit data costs bandwidth and possibly compute time. Blocking IPS is much cheaper for the blocker.
- fuckaj 1y agoZip bomb?
- aspenmayer 1y agoDoesn’t that tie up a socket on the server similarly to how a keepalive would on the bot user end?
- recursive 1y agoI don't think so. The payload size of the bytes on the wire is small. This premise is all dependent on the .zip being crawled synchronously by the same thread/job making the request.
- aspenmayer 1y agoWhat if bots catch on to zip bombs, and just download them really slowly? https://en.wikipedia.org/wiki/Zeno%27s_paradoxes#Dichotomy_paradox https://en.wikipedia.org/wiki/Zeno%27s_paradoxes#Dichotomy_p...
- throwawayffffas 1y agoTheir objective is not to DDOS websites, if they catch on, they will download it fast and then discard it.
- praptak 1y ago"I'm seriously thinking that the CCP encourage this with maybe the hope of externalizing the cost of the Great Firewall to the rest of the world. If China scrapes content, that's fine as far as the CCP goes; If it's blocked, that's fine by the CCP too (I say, as I adjust my tin foil hat)." Then turn the tables on them and make the Great Firewall do your job! Just choose a random snippet about illegal Chinese occupation of Tibet or human rights abuses of Uyghur people each time you generate a page and insert it as a breaker between paragraphs. This should get you blocked in no time :)
- blueflow 1y agoI just tried this, i took some strings about Falun Gong and the Tianmen thing from the chinese wikipedia and put them into my SSH server banner. The connection attempts from the Tencent AS ceased completely, but now they come from Russia, Lithuania and Iran instead.
- gitpusher 1y agoWhoa, that's fascinating. So their botnet runs in multiple regions and will auto-switch if one has problems. Makes sense. Seems a bit strange to use China as the primary, though. Unless of course the attacker is based in China? Of the countries you mentioned Lithuania seems a much better choice. They have excellent pipes to EU and North America, and there's no firewall to deal with
- that_lurker 1y agoWhy not just block the User Agent?
- N_Lens 1y agoBots often rotate the UA too, their entire goal is to get through and scrape as much content as possible, using any means possible.
- aspenmayer 1y agoI think the UA is easily spoofed, whereas the AS and IP are less easily spoofed. You have everything you need already to spoof UA, while you will need resources to spoof your IP, whether it’s wall clock time to set it up, CPU time to insert another network hop, and/or peers or other third parties to route your traffic, and so on. The User Agent are variables that you can easily change, no real effort or expense or third parties required.
- lexicality 1y agobecause you have to parse the http request to do that, while blocking the IP can be done at the firewall
- arewethereyeta 1y agoBecause it's the single most falsifiable piece of information you would find on ANY "how to scrape for dummies" article out there. They all start with changing your UA.
- ryantgtg 1y agoSure, but the article is about a bot that expressly identifies itself in the user agent and its user agent name contains a sentence suggesting you block its ip if you don’t like it. Since it uses at least 74 ips, blocking its user agent seems like a fine idea.
- geokon 1y agoIs there a way to reverse look up IPs by company? Like a list off all IPs owned by Alphabet, Meta Bing etc?
- BLKNSLVR 1y agohttps://hackertarget.com/as-ip-lookup/ https://hackertarget.com/as-ip-lookup/ Chuck 'Tencent' into the text box and execute.
- boris 1y agoYes, I've seen this one in our logs. Quite obnoxious, but at least it identifies itself as a bot and, at least in our case (cgit host), does not generate much traffic. The bulk of our traffic comes from bots that pretend to be real browsers and that use a large number of IP addresses (mostly from Brazil and Asia in our case). I've been playing cat and mouse trying to block them for the past week and here are a couple of observations/ideas, in case this is helpful to someone: * As mentioned above, the bulk of the traffic comes from a large number of IPs, each issuing only a few requests a day, and they pretend to be real UAs. * Most of them don't bother sending the referrer URL, but not all (some bots from Huawei Cloud do, but they currently don't generate much traffic). * The first thing I tried was to throttle bandwidth for URLs that contain id= (which on a cgit instance generate the bulk of the bot traffic). So I set the bandwidth to 1Kb/s and thought surely most of the bots will not be willing to wait for 10-20s to download the page. Surprise: they didn't care. They just waited and kept coming back. * BTW, they also used keep alive connections if ones were offered. So another thing I did was disable keep alive for the /cgit/ locations. Failed that enough bots would routinely hog up all the available connections. * My current solution is to deny requests for all URLs containing id= unless they also contain the `notbot` parameter in the query string (and which I suggest legitimate users add in the custom error message for 403). I also currently only do this if the referrer is not present but I may have to change that if the bots adapt. Overall, this helped with the load and freed up connections to legitimate users, but the bots didn't go away. They still request, get 403, but keep coming back. My conclusion from this experience is that you really only have two options: either do something ad hoc, very specific to your site (like the notbot in query string) that whoever runs the bots won't bother adapting to or you have to employ someone with enough resources (like Cloudflare) to fight them for you. Using some "standard" solution (like rate limit, Anubis, etc) is not going to work -- they have enough resources to eat up the cost and/or adapt.
- palmfacehn 1y agoPick an obscure UA substring like MSIE 3.0 or HP-UX. Preemptively 403 these User Agents, (you'll create your own list). Later in the week you can circle back and distill these 403s down to problematic ASNs. Whack moles as necessary.
- teekert 1y agoThese IP addresses being released at some point, and making their way into something else is probably the reason I never got to fully run my mailserver from my basement. These companies are just massively giving IP addresses a bad reputation, messing them up for any other use and then abandoning them. I wonder what this would look like when plotted: AI (and other toxic crawling) companies slowly consuming the IPv4 address space? Ideally we'd forced them into some corner of the IPv6 space I guess. I mean robots.txt seems not to be of any help here.
- niczem 1y agoI think banning IPs is a treadmill you never really get off of. Between cloud providers, VPNs, CGNAT, and botnets, you spend more time whack-a-moling than actually stopping abuse. What’s worked better for me is tarpitting or just confusing the hell out of scrapers so they waste their own resources. There’s a great talk on this: Defense by numbers: Making Problems for Script Kiddies and Scanner Monkeys https://www.youtube.com/watch?v=H9Kxas65f7A https://www.youtube.com/watch?v=H9Kxas65f7A What I’d really love to see - but probably never will—is companies joining forces to share data or support open projects like Common Crawl. That would raise the floor for everyone. But, you know… capitalism, so instead we all reinvent the wheel in our own silos.
- BLKNSLVR 1y agoIf you can automate the treadmill and set a timeout at which point the 'bad' IPs will go back to being 'not necessarily bad', then you're minimising the effort required. An open project that classifies and records this - would need a fair bit of on-going protection, ironically.
- adinhitlore 1y ago[flagged]
- BLKNSLVR 1y agoI've mentioned my project[0] before, and it's just as sledgehammer-subtle as this bot asks. I have a firewall that logs every incoming connection to every port. If I get a connection to a port that has nothing behind it, then I consider the IP address that sent the connection to be malicious, and I block the IP address from connecting to any actual service ports. This works for me, but I run very few things to serve very few people, so there's minimal collateral damage when 'overblocking' happens - the most common thing is that I lock myself out of my VPN (lolfacepalm). I occasionally look at the database of IP addresses and do some pivot tabling to find the most common networks and have identified a number of cough security companies that do incessant scanning of the IPv4 internet among other networks that give me the wrong vibes. [0]: Uninvited Activity: https://github.com/UninvitedActivity/UninvitedActivity https://github.com/UninvitedActivity/UninvitedActivity P.S. If there aren't any Chinese or Russian IP addresses / networks in my lists, then I probably block them outright prior to the logging.
- rglullis 1y ago[flagged]
- latexr 1y ago> All of the "blockchain is only for drug dealing and scams" people will sooner or later realize that it is the exact type of scenarios that makes it imperative to keep developing trustless systems. This is like saying “All the “sugar-sweetened beverages are bad for you” people will sooner or later realize it is imperative to drink liquids”. It is perfectly congruent to believe trustless systems are important and that the way the blockchain works is more harmful than positive. Additionally, the claim is that cryptocurrencies are used like that. Blockchains by themselves have a different set of issues and criticisms.
- rglullis 1y agoTell that to the "web3 is doing great" crowd. I've met and worked with many people who never shilled a coin in their whole life and were treated as criminals for merely proposing any type of application on Ethereum. I got tired of having people yelling online about how "we are burning the planet" and who refused to understand that proof of stake made energy consumption negligible. To this day, I have my Mastodon instance on some extreme blocklist because "admin is a crypto shill" and their main evidence was some discussion I was having to use ENS as an alternative to webfinger so that people could own their identity without relying on domain providers. The goalposts keep moving. The critics will keep finding reasons and workarounds. Lots of useful idiots will keep doubling down on the idea that some holy government will show up and enact perfect regulation, even though it's the institutions themselves who are the most corrupt and taking away their freedoms. The open, anonymous web is on the verge of extinction. We no longer can keep ignoring externalities. We will need to start designing our systems in a way where everyone will need to either pay or have some form of social proof for accessing remote services. And while this does not require any type of block chains or cryptocurrency, we certainly will need to start showing some respect to all the people who were working on them and have learned a thing or two about these problems.
- 1y ago
- mellosouls 1y agoUnfortunately, HN itself is occasionally used for publicising crawling services that rely on underhand techniques that don't seem terribly different to the ones here. I don't know if its because they operate in the service of capital rather than China, as here, but use of those methods in the former case seems to get more of a pass here.
- bob1029 1y agoI think a lot of really smart people are letting themselves get taken for a ride by the web scraping thing. Unless the bot activity is legitimately hammering your site and causing issues (not saying this isn't happening in some cases), then this mostly amounts to an ideological game of capture the flag. The difference being that you'll never find their flag. The only thing you win by playing is lost time. The best way to mitigate the load from diffuse, unidentifiable, grey area participants is to have a fast and well engineered web product. This is good news, because your actual human customers would really enjoy this too.
- phito 1y agoMy friend has a small public gitea instance, only use by him a a few friends. He's getting thousounds of requests an hour from bots. I'm sorry but even if it does not impact his service, at the very least it feels like harassment
- dmesg 1y agoYes and it makes reading your logs needlessly harder. Sometimes I find an odd password being probed, search for it on the web and find an interesting story, that a new backdoor was discovered in a commercial appliance. In that regard reading my logs led me sometimes to interesting articles about cyber security. Also log flooding may result in your journaling service truncating the log and you miss something important.
- wvbdmp 1y agoYou log passwords?
- deleted 1y ago[deleted]
- zeta0134 1y agoJust about nobody logs passwords on purpose. But really stupid IoT devices accept credentials as like query strings, or part of the path or something, and it's common to log those. The attacker is sending you passwords meant for a much less secure system.
- baxuz 1y agoGeoblocking China and Russia should be the default.
- mediumsmart 1y agofiefdom internet 1.0 release party - information near your fingertips
- Jnr 1y agoExternally I use Cloudflare proxy and internally I put Crowdsec and Modsecurity CRS middlewares in front of Traefik. After some fine-tuning and eliminating false positives, it is running smoothly. It logs all the temporarily banned and reported IPs (to Crowdsec) and logging them to a Discord channel. On average it blocks a few dozen different IPs each day. From what I see, there are far more American IPs trying to access non-public resources and attempting to exploit CVEs than there are Chinese ones. I don't really mind anyone scraping publicly accessible content and the rest is either gated by SSO or located in intranet. For me personally there is no need to block a specific country, I think that trying to block exploit or flooding attempts is a better approach.
- poisonborz 1y agoCrowdsec: the idea is tempting, but giving away all of the server's traffic to a for-profit is a huge liability.
- Jnr 1y agoYou pass all traffic through Cloudflare. You do not pass any traffic to Crowdsec, you detect locally and only report blocked IPs. And with Modsecurity CRS you don't report anything to anyone but configuring and fine tuning is a bit harder.
- jrgifford 1y agoThe more egregious attempts are likely being blocked by Cloudflare WAF / similar.
- Jnr 1y agoI don't think they are really blocking anything unless you specifically enable it. But it gives some piece of mind knowing that I could probably enable it quickly if it becomes necessary.
- harbingerofdoom 1y ago[dead]
- poisonborz 1y agoNaive question: why isn't there a publicly accessible central repository of bad IPs and domains, stewarded by the industry, operated by a nonprofit, like W3C? Yes it wouldn't be enough by itself ("bad" is a very subjective term) but it could be a popular well-maintained baseline.
- sim7c00 1y agothere are many of these and they are always outdated another issue is things like cloud hosting will overlap their ranges with legit business ranges happily, so if you go that route you will inadvertently also block legitimate things. not that a regular person care too much for that, but an abuse list should be accurate.
- nubinetwork 1y agohttps://xkcd.com/927/ https://xkcd.com/927/ For what it's worth, I'm also guilty of this, even if I made my site to replace one that died.
- timpera 1y agoAs someone who uses VPNs all the time, these comments make me sad. Blocking by IP is not the solution.
- deleted 1y ago[deleted]
- sim7c00 1y agoi think there is an opportunity to train an neural network on browser user agent s(they are catalogued but vary and change a lot). then u can block everything not matching. it will work better than regex. a lot of these companies rely on 'but we are clearly recognizable' via fornexample these user agents, as excuse to put burden on sysadmins to maintains blocklists instead of otherway round (keep list of scrapables..) maybe someone mathy can unburden them ? you could also look who ask for nonexisting resources, and block anyone who asks for more than X (large enough not to let config issue or so kill regular clients). block might be just a minute so u dont have too many risk when an FP occurs. it will be enough likely to make the scraper turn away. there are many things to do depending on context, app complexity, load etc. , problem is there's no really easy way to do these things. ML should be able to help a lot in such a space??
- arewethereyeta 1y agoWhat exactly do you want to train on a falsifiable piece of info? We do something like this at https://visitorquery.com https://visitorquery.com in order to detect HTTP proxies and VPNs but the UA is very unreliable. I guess you could detect based on multiple pieces with UA being one of them where one UA must have x, y, z or where x cannot be found on one UA. Most of the info is generated tho.
- ricardo81 1y agoAny respectable web scale crawler(/scraper) should have reverse DNS so that it can automatically be blocked. Though it would seem all bets are off and anyone will scrape anything. Now we're left with middlemen like cloudflare that cost people millions of hours of time ticking boxes to prove they're human beings.
- breve 1y ago> Alex Schroeder's Butlerian Jihad That's Frank Herbert's Butlerian Jihad.
- dirkc 1y agoInteresting to think that the answer to banning thinking computers in Dune was basically to indoctrinate kids from birth (mentats) and/or doing large quantities of drugs (guild navigators).
- flanbiscuit 1y agoTo be fair, he was referring to a post on Alex Schroeder's blog titled with the same name as the term from the Dune books. And that post correctly credits Dune/Herbert. But the post is not about Dune, it's about Spam bots so it's more related to what the original author's post is about. Speaking of the Butlerian Jihad, Frank Herbert's son (Brian) and another author named Kevin J Anderson co-wrote a few books in the Dune universe and one of them was about the Butlerian Jihad. I read it. It was good, not as good at Frank Herbert's books but I still enjoyed it. One of the authors is not as good as the other because you can kind of tell the writing quality changing per chapter. https://en.wikipedia.org/wiki/Dune:_The_Butlerian_Jihad https://en.wikipedia.org/wiki/Dune:_The_Butlerian_Jihad
- renewiltord 1y agoThat's really hard to believe. Brian Herbert's stuff seems sort of Fan Fiction In The World of Dune. Nothing wrong with fan fiction: The Last Ringbearer etc. are pretty enjoyable. But BH just follows on. His work has a bit of the feeling of people who lived in the ruins of the roman forum https://x.com/museiincomune/status/1799039086906474572 https://x.com/museiincomune/status/1799039086906474572
- int_19h 1y agoThose books completely misrepresent Frank Herbert's original ideas for Butlerian Jihad. It wasn't supposed to be a literal war against genocidal robots.
- hackrmn 1y agoI know opinions are divided on what I am about to mention, but what about CAPTCHA to filter bots? Yes, I am well aware we're a decade past a lot of CAPTCHA being broken by _algorithms_, but I believe it is still a relatively useful general solution, technically -- question is, would we want to filter non-humans, effectively? I am myself on the fence about this, big fan of what HTTP allows us to do, and I mean specifically computer-to-computer (automation/bots/etc) HTTP clients. But with the geopolitical landscape of today, where Internet has become a tug of war (sometimes literally), maybe Butlerian Jihad was onto something? China and Russia are blatantly and near-openly shoving their fingers in every hole they can find, and if this is normalized so will Europe and U.S., for countermeasure (at least one could imagine it being the case). One could also allow bots -- clients unable to solve CAPTCHA -- access to very simplified, distilled and _reduced_ content, to give them the minimal goodwill to "index" and "crawl" for ostensibly "good" purposes.
- Hizonner 1y agoIt's a simple translation error. They really meant "Feed me worthless synthetic shit at the highest rate you feel comfortable with. It's also OK to tarpit me."
- 8organicbits 1y agoI've been working on a web crawler and have been trying to make it as friendly as possible. Strictly checking robots.txt, crawling slowly, clear identification in the User Agent string, single IP source address. But I've noticed some anti-bot tricks getting applied to the robot.txt file itself. The latest was a slow loris approach where it takes forever for robots.txt to download. I accidentally treated this as a 404, which then meant I continued to crawl that site. I had to change the code so a robots.txt timeout is treated like a Disallow /. It feels odd because I find I'm writing code to detect anti-bot tools even though I'm trying my best to follow conventions.
- navane 1y agoThat's like deterring burglars by hiding your doorbell
- brianwawok 1y agoI doubt that’s on purpose. The bad guys that don’t follow robots don’t bother downloading it. Never assume malice what can be attributed to incompetence.
- NegativeK 1y agoI really appreciate you giving a shit. Not sarcastically -- it seems like you're actually doing everything right, and it makes a difference. Gating robots.txt might be a mistake, but it also might be a quick way to deal with crawlers who mine robots.txt for pages that are more interesting. It's also a page that's never visited by humans. So if you make it a tarpit, you both refuse to give the bot more information and slow it down. It's crap that it's affecting your work, but a website owner isn't likely to care about the distinction when they're pissed off at having to deal with bad actors that they should never have to care about.
- jedisct1 1y agoI don’t understand why people want to block bots, especially from a major player like Tencent, while at the same time doing everything they can to be indexed by Google
- orochimaaru 1y agoIs there a list of chinese ASN’s that you can ban if you don’t do much business there - eg all of China, Macau, select Chinese clouds in SE Asia, Polynesia and Africa. I think they’ve kept HK clean so far.
- zImPatrick 1y agoI run an instance of a downloader tool and had lots of chinese IPs mass-download youtube videos with the most generic UA. I started with „just“ blocking their ASNs, but they always came back with another one until I just decided to stop bothering and banned China entirely. I‘m confused on why some chinese ISPs have so many different ASNs - while most major internet providers here have exactly one.
- alphazard 1y agoI'm always a little surprised to see how many people take robots.txt seriously on HN. It's nice to see so many folks with good intentions. However, it's obviously not a real solution. It depends on people knowing about it, and adding the complexity of checking it to their crawler. Are there other more serious solutions? It seems like we've heard about "micropayments" and "a big merkle tree of real people" type solutions forever and they've never materialized.
- ralferoo 1y ago> It depends on people knowing about it, and adding the complexity of checking it to their crawler. I can't believe any bot writer doesn't know about robots.txt. They're just so self-obsessed and can't comprehend why the rules should apply to them, because obviously their project is special and it's just everyone else's bot that causes trouble.
- Bender 1y ago(malicious) Bot writers have exactly zero concern for robots.txt. Most bots are malicious. Most bots don't set most of the TCP/IP flags. Their only concern is speed. I block about 99% of port scanning bots by simply dropping any TCP SYN packet that is missing MSS or uses a strange value. The most popular port scanning tool is masscan which does not set MSS and some of the malicious user-agents also set some odd MSS values if they even set it at all. -A PREROUTING -i eth0 -p tcp -m tcp -d $INTERNET_IP --syn -m tcpmss ! --mss 1280:1460 -j DROP Example rule from the netfilter raw table. This will not help against headless chrome. The reason this is useful is that many bots first scan for port 443 then try to enumerate it. The bots that look up domain names to scan will still try and many of those come from new certs being created in LetsEncrypt. That is one of the reasons I use the DNS method, get a wildcard and sit on it for a while. Another thing that helps is setting a default host in ones load balancer or web server that serves up a default simple static page served from a ram disk that say something like, "It Worked!" and disable logging for that default site. In HAProxy one should look up the option "strict-sni". Very old API clients can get blocked if they do not support SNI but along that line most bots are really old unsupported code that the botter could not update if their life depended on it.
- ralferoo 1y agoOut of spite, I'd ignore their request to filter by IP (who knows what their intent is by saying that - maybe they're connecting from VPNs or tor exit nodes to cause disruption etc), but instead filter by matching for that content in the User-Agent instead and feeding them a zip bomb instead.
- nirui 1y ago> A further check showed that all the network blocks are owned by one organization—Tencent. I'm seriously thinking that the CCP encourage this with maybe the hope of externalizing the cost of the Great Firewall to the rest of the world. A simple check against the IP address 170.106.176.0, 150.109.96.0, 129.226.160.0, 49.51.166.0 and 43.135.0.0 showed that these IP addresses is allocated to Tencent Cloud, a Google Cloud-like rental service. I'm using their product personally, it's really cheap, a little more than $12~$20 a year for a VPS, and it's from one of the top Internet company. Sure, it can't really completely rule out the possibility that Tencent is behind all of this, but I don't really think the communist needs to attack your website through Tencent, it's just simply not logical. More likely it's just some company rented some server on Tencent crawling the Internet. The rest is probably just your xenophobia fueled paranoia.
- 1000units 1y agoThis seems like a plot to redirect residential Chinese traffic through VPNs, which are supposedly mostly operated by only a few entities with a stomach for dishonest maneuvering and surveillance.
- 1000units 1y agoMaybe they mean well. You'd have to understand their total security posture.
- nojs 1y ago> Here's how it identifies itself: “Mozilla/5.0 (compatible; Thinkbot/0.5.8; +In_the_test_phase,_if_the_Thinkbot_brings_you_trouble,_please_block_its_IP_address._Thank_you.)”. I mean you could just ban the user agent? The real issue is with bots pretending not to be bots.
- ryantgtg 1y agoThis one was funny because I checked a day’s log and it was using at least 15 different IPs. Much easier to just ban or rate limit “Thinkbot”
- xyst 1y agoIn this day and age of crab barreling over one another, simple gestures such as honoring _robots.txt_ are just completely ignored.
- johnklos 1y agoThis is a problem. There's a recent phishing campaign with sites hosted by Cloudflare and spam sent through either "noobtech.in" (103.173.40.0/24) or through "worldhost.group" (many, many networks). "noobtech.in" has no web site, can't accept abuse complaints (their email has spam filters), and they don't respond at all to email asking them for better communication methods. The phishing domains have "mail.(phishing domain)" which resolves back to 103.173.40.0/24. Their upstream is a Russian network that doesn't respond to anything. It's 100% clear that this network is only used for phishing and spam. It's trivial to block "noobtech.in". "worldhost.group", though, is a huge hosting conglomerate that owns many, many hosting companies and many, many networks spread across many ASNs. They do not respond to any attempts to communicate with them, but since their web site redirects to "hosting.com", I've sent abuse complaints to them. "hosting.com" has autoresponders saying they'll get back to me, but so far not a single ticket has been answered with anything but the initial autoresponder. It's really, really difficult to imagine how one would block them, and also difficult to imagine what kind of collateral impact that'd have. These huge providers, Tencent included, get away with way too much. You can't communicate with them, they don't give the slightest shit about harmful, abusive and/or illegal behavior from their networks, and we have no easy way to simply block them. I think we, collectively, need to start coming up with things we can do that would make their lives difficult enough for them to take notice. Should we have a public listing of all netblocks that belong to such companies and, as an example, we could choose to autorespond to all email from "worldhost.group" and redirect all web browsing from Tencent so we can tell people that their ISP is malicious? I don't know what the solution is, but I'd love to feel a bit less like I have no recourse when it comes to these huge mega-corporations.
- deleted 1y ago[deleted]
- dom-whg 1y agoCould you drop a message to dom@ with more details and I'll get this stopped from the WHG side - and find out what happened. Thanks!
- Avamander 1y agoIf you block them and they're legitimate, they'll surely find a way to actually start a dialogue. If that feels too harsh you could also start serving captchas and tarpits, but I'm unsure if it's worth actually bothering with.
- dizlexic 1y agoI've written a decent number of malicious crawlers in my time. Be happy they game you a user agent.
- thatoneguy 1y agoIs it fair game to just return fire and sink the crawlers? A few thousand connections via a few dozen residential proxies might do it.
- BizarroLand 1y agoI feel like one annoying solution to bots would be putting your pages behind a simple account creation & logon screen. Maybe it could be for your archive files or something. Still a hassle but if 95% of your blog requires a login to view that would decrease the load quite a bit, right?
- 1970-01-01 1y agoWhat about IPv6?
- bit1993 1y agoI have the client send a custom header with every request, and block all other request.
- renewiltord 1y agoI would never have considered this, but someone on HN pointed out that web user agents work like this. Servers send ads and there is no way for them to enforce that browsers render the ads and hide the content or whatever. The user agent is supposed to act for the user. "Your business model is not my problem", etc. Well, my user agents work for me, not for you - the server guy who is complaining about this and that. "Your business model is not my problem". Block me if you don't want me. https://news.ycombinator.com/item?id=44975697 https://news.ycombinator.com/item?id=44975697
- johneth 1y agoWell done on pointing out exactly what everyone here is saying in the most arrogant way possible. Also, well done on linking to your own comment where people explain this to you. The problem is that there is no way to "block me if you don't want me". That's the entire issue. The methods these scrapers use mean it's nigh on impossible to block them.
- Avamander 1y agoSo far it's actually not. Though it is getting harder. I suspect we'll get integrity attestation or tokens before it becomes an unsurmountable problem to block bots.
- renewiltord 1y ago"Your inability to engineer is not my problem". See: https://news.ycombinator.com/item?id=45018660 https://news.ycombinator.com/item?id=45018660
- reincoder 1y agoI work for IPinfo. If you need IP Address CIDR blocks for any country or ASN, let me know. I have our data in front of me and can send it over via Github Gist. Thank you.
- imoverclocked 1y agoI’ve been having a heck of a time figuring out where some malicious traffic is coming from. Nobody has been able to give me a straight answer when I give them the ip: 127.1.5.12 Maybe you can help trace-a-route to them? I’d just love to know whois behind that IP. If nothing else, I could let them know to be standards compliant and implement rfc 3514.
- genewitch 1y agothe malicious traffic is coming from inside the house
- kldg 1y agoWhat is the commonality between websites severely affected by bots? I run web server from home for years on .com TLD, is high-ish in Google site index for relevant keywords, and do not have any exotic protections against bots either on router or server (though I did make an attempt at counting bots, for curiosity). I get very frequent port scans, and they usually grab the index page, but only rarely follow dynamically-loaded links. I don't even really think about bots because there is no noticeable impact either when I ran server on Apache 2, and now with multiple websites run using Axum. I would guess directory listing? -But I'm an idiot, so any elucidation would be appreciated.
- gucci-on-fleek 1y agoFor my personal site, I let the bots do whatever they want—it's a static site with like 12 pages, so they'd essentially need to saturate the (gigabit) network before causing me any problems. On the other hand, I had to deploy Anubis for the SVN web interface for tug.org. SVN is way slower than Git (most pages take 5 seconds to load), and the server didn't even have basic caching enabled, but before last year, there weren't any issues. But starting early this year, the bots started scraping every revision, and since the repo is 20+ years old and has 300k files, there are a lot of pages to scrape. This was overloading the entire server, making every other service hosted there unusable. I tried adding caching and blocking some bad ASNs, but Anubis was (unfortunately) the only solution that seems to have worked. So, I think that the main commonality is popular-ish sites with lots of pages that are computationally-expensive to generate.
- inetknght 1y agoIf a US IP is abusing the internet, you can go through the courts. If a foreign country is... good luck. So, are hackers and internet shittery coming from China? Block China's ASNs. Too bad ISPs won't do that, so you have to do it yourself. Keep it blocked until China enforces computer fraud and abuse.
- socalgal2 1y agoAren't many apartment buildings all coming from just a few IP addresses? https://en.wikipedia.org/wiki/Carrier-grade_NAT https://en.wikipedia.org/wiki/Carrier-grade_NAT
- TZubiri 1y agoYes, and this makes ip banning have false positives. But ultimately it's worth it, you are responsible for your neighbours.
- simoncion 1y ago> [Y]ou are responsible for [how] your neighbours [use the Internet]. Nope. I'm very much not responsible for snooping on my neighbor's private communications. If anyone is responsible for doing any sort of abuse monitoring, it is the ISP chosen by my neighbor.
- TZubiri 1y agoThis is not a normative social prescription, but a descriptive natural phenomenon. If there's a neighbour in your building who is running a bitcoin farm on your residential building, it's going to cause issues for you. If people from your country commit crime in other countries and violate visas, then you are going to face a quota due to them. If you bank at ACME Bank, and then it turns out they were arms traffickers, your funds were pooled and helped launder their money, you are responsible by association . Reputation is not only individual, but there is group reputation, regardless of whether you like it or not.
- simoncion 1y ago> If there's a neighbour... What ass-backwards jurisdiction do you live in where any of the things you mention in this paragraph are true, let alone the notion that uninvolved bystanders would be responsible for the behavior of others?
- 1y ago
- TZubiri 1y agoIs there an argument to be made that circumventing bans by changing ip addresses is illegal? Businesses have right of refusal!
- topak3000 1y agoYou can block traffic by AS effectively. In my case, I have seen a large number of crawlers from Tencent and Alibaba. I signed up for the free AS database from IP2Location LITE and then blocked the ranges from those ASNs.
- loph 1y agothis is related: https://www.pcworld.com/article/2845330/hundreds-of-chrome-extensions-create-a-web-scraping-botnet.html https://www.pcworld.com/article/2845330/hundreds-of-chrome-e... The TL;DR is that there are malicious browser plugins that make the browser into a web scraping bot. I see this all the time in web server logs; it is recognizable as a GET on a deep link coming from some random IP, usually residential.