11 ms·
An update on residential proxies and the scraper situation
- cyanydeez 3mo agommm, in many cases these residential proxies are media boxes, and they consent as much as anyone else consents to what amazon, or google or facebook does; it's buried somewhere in the recesses of the TOS. The question is more about why the US and others can't properly enforce the bullshit all this amounts to.
- bell-cot 3mo ago"He who has the gold makes the rules" is older than the pyramids.
- SR2Z 3mo agoBecause this isn't clearly against the law, nor should it be. If websites want to ban based on IP address lots of innocent users get caught in the cross-fire. I'm not sure what the solution would look like - maybe Cloudflare's payment required for requests beyond a certain limit? But I think that the world needs user freedoms now more than ever.
- mschuster91 3mo ago> The question is more about why the US and others can't properly enforce the bullshit all this amounts to. It would cost too much money, either for police to raid all the physical shops and ebay sellers selling dodgy IPTV boxes, or for ISPs to hire enough competent support staff to monitor and respond to abuse@ email addresses and follow through.
- TurdF3rguson 3mo agoWhat exactly should be illegal here? Scraping websites? AI agents? Not following robots.txt?
- aorth 3mo agoThe excessive scraping and ignoring robots.txt only breaks the informal social contract established over the past decades of the open internet. The real problem is the companies offering money to developers if they include unrelated SDKs in their calculator or flashlight (for example) applications. Those SDKs add functionality to incorporate those devices into a network that can be used for scraping. The traffic is little, but is distributed over millions of residential devices all over the world, making it difficult to categorize or block. That should be illegal, and that's what Google et al can be expected to be policing on their app stores.
- inigyou 3mo agoOn what grounds would it be illegal though? Things don't become illegal just because you don't like them. They may become illegal just because the president doesn't like them, but I don't think you're him, and in the absence of that, there has to be a majority of Congress and most of them want a reason.
- robinsonb5 3mo agoBecause they don't have the informed consent* of the owner of the device wich ends up running the code? * no, small print in a click-through agreement doesn't count.
- inigyou 3mo agoIt's not illegal to run code on a device without informed consent to everything the code does. The CFAA may be excessively broad but it isn't that broad.
- TurdF3rguson 3mo agoI think they do actually. It's pretty clear from the consent screen that they're doing what they're doing.
- 3mo ago
- tingletech 3mo agoThe comments are not showing up for me now, but when they were still showing for anonymous users, there was a link to https://commoncrawl.org https://commoncrawl.org. I've been sort of worried about letting agents hit websites, I wonder if a fetch_url agent tool could be made to look in common crawl first before hitting the web for it?
- colinsane 3mo agojust their smallest dataset looks to be 6 TB _compressed_. not a thing you can really ship as part of the agent. but if somebody made a fetch_url tool that sharded that across all users of it, i'd give it a try. could probably just layer that on top of bittorrent or IPFS or something.
- atomic128 3mo agoThere is a large community of people that poison scrapers. The poison gets better every day, and the community is continuously growing. Poison Fountain, alone, transmits hundreds of gigabytes of poison per day, which goes into scrapers, git repositories on every hosting platform, social media, etc. Part of the poisoning community on Reddit, for example: https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_new_version_of_poison_fountain_is_up_and/ https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_n...
- logancbrown 3mo agoPeople think this is causing issues for data collection for LLMs, but in reality it's not and there are several very trivial mechanisms to employ in data collection to bypass the "poison data" issue. The internet landscape was already poisoned with fake data, fringe conspiracies, and text before this Poison Fountain initiative.
- zuzululu 3mo agoexactly i took a look at that subreddit and doesnt look like theres any professionals just bunch of anti-AI users who thinks they are smarter its very easy to detect and bypass poison type of tools largely because of the fact that there are far more outlets for truthful info so unless you can get everyone to buy in (with real legal liabilities) its not effective also its possible to poison the poisoners with a certain pill that would have very real consequences for those maintaining whatever github repo/communities
- andai 3mo agoYeah. A fun thing to do is to try and actually read common crawl! Really makes you think, what we're feeding them...
- dang 3mo agoI've banned this account because we don't allow single-purpose accounts on HN, and your account has been doing that for quite some time now. We ban such accounts regardless of what the single purpose happens to be. Pre-existing agendas are not what HN is for and destroy the curious conversation that it is supposed to be for. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html Edit: If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html.
- mips_avatar 3mo agoI feel like the solution is a better common crawl. As nice as it would be to block the frontier AI labs from getting access to information, we should reset the baseline of information accessibility so there's less marginal advantage on these labs. I worry a lot of the anti scraping rhetoric will just injure the open web and put somebody like cloudflare in charge.
- jay_kyburz 3mo agoI agree, if up-to-data data was available somewhere else and free, there would be no reason to pay hackers and scrape. You could perhaps even get website operators to "push" new data to a common crawl database. The scrapers would learn there is no value on scraping X domain because the data is available elsewhere more easily.
- jay_kyburz 3mo agoHow about a website header with a link to a static zip that contains the whole website in one hit. The Zip could be hosted on some big public sever. Perhaps even mirrored locally for each nation.
- mips_avatar 3mo agothat's hard to do with rendered content, oftentimes the result depends on a backend service. Maybe you should make the service it's running public but that might be a line most aren't willing to cross.
- jay_kyburz 3mo agoI was thinking you scrape your own website every day in the middle of the night when traffic is low, and make that available. They can come and collect it every day if they want to.
- mips_avatar 3mo ago
- everfrustrated 3mo agoI wonder how much of this is traffic caused by peoples agents using web tools causing searches and fetches rather than general trawls of the internet.
- corbet 3mo agoVery little of it. When you see a million IPs systematically working their way through your URL space, it's pretty clear that there's a central control node behind it all.
- everfrustrated 3mo agoYour earlier article suggests you aren't using a CDN. Might be well worth looking into - not for any bot detection so much as just having a good old fashioned cache in front of you.
- mplewis 3mo agoAs someone who operates a wiki, this does not solve the problem.
- solid_fuel 3mo agoCaches only help for pages that have been requested recently. The behavior of crawlers - going from one page to the next across the whole site - will probably not be mitigated significantly by a cache.
- dawnerd 3mo agoI've seen some logs where a bunch of random ips were hitting a client's search endpoint feeding what looked like user questions to it. Of course none of them returned anything useful but it was causing a lot of strain and even causing the site to go down (gotta love wordpress's stock search). I'm guessing the training companies are taking real/synthesized user queries and trying to distill what they can from site searches.
- noxvilleza 3mo ago
- sixtyj 3mo agoThe issue with scrapping is the intensity and volume of bots. I think that nobody would care if I use wget or curl for few pages, e.g. because I would like to read a site as offline or archive it. Btw average age of any page is 10 years. Deletion or structural change after acquisition is common, Signal vs Noise site recent wipe out could serve as an example why we need to archive sites.
- ccgreg 3mo agoA lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CCBot.
- sumedh 3mo agoCloudflare can verify certain bots when they come from known ip addresses. So if your site is using cloudflare it can let CCBot if it has done the verification.
- m463 3mo agocloudflare routinely denies my human-piloted browser now, on many sites.
- ambigious7777 3mo agoyou're not a bot thats irrelevant
- pocksuppet 3mo agoI wonder if you'd have more reliable internet browsing with a browser that was a CF verified bot.
- pineapplepizza6 3mo ago[dead]
- Bratmon 3mo agoResidential Proxies are the most emblematic technology of our era- a group of people looked at something that used to be considered a crime (botnets) and realized that if they just did it openly, no one would ever punish them.
- BoorishBears 3mo agoThank god for residential proxies. Highly unethical but the way the internet is going they're the last anti-hero of a somewhat open internet
- zuzululu 3mo agoi know a few very large startups that used it to fake their way into an exit unethical yes but really raises the question as to what we see is real or not
- morkalork 3mo agoMoney is real. DAU that don't pay subscriptions, or don't lead to paid conversions on hosted ads, are worthless.
- BoorishBears 3mo ago"Raises the question of what we see is real" No they really don't, dishonest founders do that. You're one with the lower case shibboleth so I have no doubt you surround yourself with dishonest founders, but faking users is pretty damn low on the usecases for residential proxies. I said they're unethical because they tend to be hidden in innocuous seeming apps or sprung on unwitting individuals via clickwraps on their smart devices.
- zuzululu 3mo agoive seen unscrupulous founders fake traction during diligence, which is my day job but ive never seen one raise $4.5m for an ai agent startup built around pulling fresh web data, then openly cheer the unethical proxy infrastructure used to evade consent and blocks then inventing a fantasy about who i associate with instead of answering that conflict is an unusually loud form of projection
- stefantalpalaru 3mo ago[dead]
- eduction 3mo agoCan BitTorrent’s architecture contribute anything useful here? I admit this is a naive question. I have no idea how applicable bt is to web requests. This problem just seems to have a similar “too many people want this resource” shape.
- fragmede 3mo agoYes but it's getting bot owners to use it is the problem. There's already the common crawl repository to start with but it isn't being used.
- dylan604 3mo agoas well as the bot owners could would never believe that the torrent has been kept up to date. the only way to do that would compare to the actual site, so why not just scrape the actual site and be done with it?
- ccgreg 3mo agoCommon Crawl's archive has metadata that says when each record (html file) was crawled.
- charcircuit 3mo agoBut who stores the metadata for the last date the site updated so you know if it needs to be refetched or not.
- ccgreg 3mo agoWe do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a page we've crawled before.
- charcircuit 3mo ago
- tiahura 3mo agoAgain, why do we allow China on the Internet? Backbone operators should not be allowed to knowingly maintain connections to networks that allow connections from China or Russia.
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- dang 3mo agoOne article mentioned in the OP was discussed here: Disrupting the largest residential proxy network - https://news.ycombinator.com/item?id=46802748 https://news.ycombinator.com/item?id=46802748 - Jan 2026 (221 comments)
- fragmede 3mo agoHow does HN fare with scraper load? Is it just CDN and pay the extra bandwidth bill for anon hit requests?
- dredmorbius 3mo agoFor one datapoint ... I have a custom HN CSS which includes some formatting of different sets of user accounts. Admins, for example, get orange highlighting and a dragon emoji (for one does not meddle in the affairs of ...). Also included are leaders, which is the one part of my CSS build script which is, or at least was until a few minutes ago, dynamic. Presently HN is returning "sorry" to my curl request. Given that I run that build manually a few times a month, it's not a matter of hitting HN with frequent scrapes. But HN has become increasingly scrape-hostile over time. Back in 2023 I did a crawl of all of HN's front-page daily history (365.25 days/year * 17 years, so about 6,200 requests), to answer a question which had come up about what was/wasn't mentioned in submission titles. That scrape included a delay (probably either 1 or 10 seconds, possibly more, I don't recall which and may have run the fetch directly from the command line), and ran (initially) without issues. I don't think it would fly today. I reported on findings at the time and several times since: <https://news.ycombinator.com/item?id=36078578 https://news.ycombinator.com/item?id=36078578> <https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=by%3Adredmorbius%20front%20page%20analysis&sort=byPopularity&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...>
- fragmede 3mo agoHN is exported to firebase, which you can hit directly, for that sort of purpose https://github.com/HackerNews/API https://github.com/HackerNews/API
- deleted 3mo ago[deleted]
- andai 3mo ago>There are ways to tell the difference — the bots usually do not fetch images or CSS, for example — but, by the time that determination is made, the address in question will not be used again. Blocking the address at that point is just a waste of time. I don't get it. Don't we keep blacklists of this stuff? And if they hammer thousands of requests per site per second and never reuse an IP, they'd run out of addresses in a few weeks. Then they'd switch to IPv6, and... well, are we using IPv6 for anything important? Like we need it for IoT, but do you want random IoT devices talking to your web server? (IPv4 handled mobile phones just fine not that long ago, right?)
- dylan604 3mo ago> do you want random IoT devices talking to your web server? Probably not, but since IoT manufacturers did zero to lock down their devices, those devices are doing a lot more than their owners think they are doing
- andai 3mo agoYes that was my point. We should just block all of them.
- Aachen 3mo agoBlocking in ipv6 works roughly the same way as in ipv4, just that the scale is different. Instead of blocking something like a company's /24 or an ISP's /16 when they don't respond to abuse messages, you block the company's /48 or the ISP's /32. It'll vary per organisation how large a range they got exactly but you can see that in WHOIS. End users are no longer at a /32 (v4) but at /64 (v6), or some prosumers might have a /29 (v4) and /56 (v6). Same concept, just a different prefix length
- inigyou 3mo agoVarious ISPs give out either a /64, a /56 or a /48. Anything else is very unusual. A normal approach is to limit by /64 at first, limit a /56 to 3-4 times the rate limit of a /64 and a /48 to 3-4 times that again. Someone who has a /48 gets to enjoy 16 times the rate limit of someone who has a /64 but that's not too bad, and you don't have to tune anything per ISP. Anyone who has bigger than /48 is no longer an individual user. They could be an ISP set up solely for scraping, though that comes with some fees and requirements.
- zb3 3mo ago> widespread scraping of web sites in search of training data for large language models and related projects This is a good thing, thanks to this we have powerful open source LLMs. > This activity overwhelms sites with traffic. When LLMs get good enough, we won't need those sites anymore :) [not satire, this is what I think, without self-censorship]
- arjie 3mo agoWhat a pity. Mostly I just want personal archives of things so that I can search them much faster than commercial solutions and the like.
- harshreality 3mo ago> ...we have tried to minimize the impact on real readers as much as possible. We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when you have millions of other people's machines to do the work on. It's massively less annoying than a captcha, which is both a longer delay (typically, at present) and a massive cognitive distraction/roadblock. The anubis author has stated they recognize it's an arms race, but PoW scales. Captchas and other signals are already at the end of the road; any additional difficulty increases false bot-positives, which are already unacceptably high. For websites running dynamic languages, a binary (anubis is in go) sentry that operates before[1] the website is forced to expend any resources, is usually a large improvement over a site-hosted captcha. I would rather, and I think most humans would agree, have to wait a few seconds, maybe even closer to a minute in the future, to get a website access token good for a day or a week, than be forced to solve a captcha. The dilemma for bots: when tokens are bound to the connecting ip, scrapers must limit the connecting IP pool for each site they want to scrape, becoming much more obvious and easy to block, or they have to use massive amounts of compute. [1] this is true regardless of whether anubis is in reverse proxy mode or auth mode.
- Groxx 3mo agoAnubis is by far the least annoying throttler I encounter. Entirely agreed, just crank it up when you get a flood, I much prefer waiting a couple seconds to interacting with custom UI for tens of seconds. I'm so glad to see that (essentially) HashCash is coming back. Now we just need it for email, like it was originally designed for...
- deleted 3mo ago[deleted]
- alightsoul 3mo agoFrom my understanding this is also how cloudflare bot protection has worked for a long time, and then they look for entropy in user input to confirm the user is human. Also how recaptcha without images works.
- 627467 3mo agoHas no one noticed their miniflux instance failing to fetch feeds because of this?
- aendruk 3mo agoSo far I’ve only encountered one site that blocked automated access to its web feed. I assume that was just an oversight.
- WarOnPrivacy 3mo agohttps://archive.fo/PAcF5 https://archive.fo/PAcF5
- rao-v 3mo agoI’m skeptical that the problem they are trying to solve is truly unreasonable bandwidth demands. Sometimes it feels like what people want is to only serve websites and content to good normal users but not evil bad “scrapers” (because maybe maybe your content will be monetized in some nebulous way) but … you put your content up publicly on the web! That should be part of reasonable use! EDIT: Lwn.net is perhaps not a fair target of my ire. “There is also a desire to not impede the operation of legitimate search engines, the Internet Archive, and other such groups. Some sites may add explicit allowlists to, for example, give the dominant search engine access to the site. Such measures have the effect of further entrenching a monopoly that already serves us poorly and should be avoided. We have, thus far, succeeded in that.” Is reasonable! Many others are not
- Catloafdev 3mo agoIf it weren't a real problem, these types of articles and services wouldn't exist.
- inigyou 3mo agoPlenty of complaints exist about things that are not real problems.
- Catloafdev 3mo agoWell that's not what's happening here, lol.
- TZubiri 3mo agoWhy the air quotes? Evading a ban and using (potentially ill gotten) residential ips to circumvent that refusal of service, is a bad actor.
- inigyou 3mo agoSurely that depends on the motivation for the ban.
- TZubiri 3mo ago>We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when you have millions of other people's machines to do the work on. The first argument that it introduces delays to users is solid, but I would advise reconsidering on the second one that a PoW workaround will be found. The moment it does you'll be able to tell because Bitcoin will crash to 0. Will bots use infected computers to do compute to work around it? Maybe, but it requires a CPU in addition to a network reputation, 2 mechanisms are stronger than one.
- inigyou 3mo agoResidential proxy users don't have the ability to run compute on their proxies.
- TZubiri 3mo agohttps://salad.com/salad-gateway-service https://salad.com/salad-gateway-service https://salad.com/earn https://salad.com/earn
- CodesInChaos 3mo ago> The moment it does you'll be able to tell because Bitcoin will crash to 0. The "workaround" for PoW is running the PoW computation on hardware that's better suited for the task. Bitcoin mining has been using ASIC for many years now. Let's say a legitimate user is willing to wait for one minute on a budget phone. Then your PoW is limited to what that phone can compute in one minute. But on the attacker's specialized hardware this computation only costs fractions of a penny, so they are barely hindered by it. The SHA256 based PoW scheme has a very heavy ASIC advantage. People have tried to design PoW scheme that minimize the custom hardware advantage, but I'm not sure if they managed to close the gap far enough to make PoW feasible for this application.
- TZubiri 3mo ago
- Avery29 3mo agoGoogle itself is a huge database.Who makes these rules depends on who's leading the market.
- ValentineC 3mo agoFrom the article: > More recently, media-streaming devices have been identified as a major carrier of malicious scraping software. Sometimes the devices are compromised at the source; other times, they are just poorly secured and easily compromised after the fact. I run an OPNsense firewall at home and the OpenWRT router at a hackerspace. Are there ways of auditing that devices aren't compromised? Tracking which devices still send lots of data when no one else is using the network?
- gucci-on-fleek 3mo ago> Tracking which devices still send lots of data when no one else is using the network? That's what I personally do at least: I have nlbwmon [0] installed on my OpenWRT router to track data usage per device, then I scrape it every minute with Prometheus and plot it in Grafana [1]. This helps me see if any IoT devices are compromised, but it probably won't help much if people are using sketchy free VPNs on their phones. I also adblocking enabled on my router [2], which helps block a few malicious domains (but certainly isn't a panacea). [0]: https://github.com/jow-/nlbwmon https://github.com/jow-/nlbwmon [1]: https://www.maxchernoff.ca/files/grafana-network-bandwidth.png https://www.maxchernoff.ca/files/grafana-network-bandwidth.p... [2]: https://docs.mossdef.org/adblock-fast/ https://docs.mossdef.org/adblock-fast/
- Aachen 3mo agoOpnsense has a traffic capture feature in the interface diagnostics menu, if you want to spot check what servers the devices are currently talking to. Should be pretty obvious: client devices and internal services will have no traffic >95% of the time, just NTP for timekeeping, DHCP lease renewal, and associated ARP (running total: two dozen packets if you monitor them for a full 24h), then any system updaters (readily identifiable by the initial DNS requests), and finally of course you'll see the traffic of the service that the device hosts, if any, which can be easily dismissed by not looking at incoming connections (scraping uses outgoing connections)
- charcircuit 3mo ago>types of operator running residential-proxy networks to attack web sites. This is such a malicious interpretation. Do you think VPN operating are also trying to attack websites? Both offer the same kind of product. >paid for hijacking their users' network connections Nothing is being hijacked. Again the author is using wording to try and paint these people as malicious actors. >Recently, LWN was subjected what was, by far, the heaviest scraper attack yet. LWN is a static site. To me it seems more expensive to use Anubis than just serve the actual page. >will now check for NetNut-infected apps Apps are not infected with NetNut. This is just Google abusing their monopoly position to hurt its competitors.
- throwaway7356 3mo ago> Apps are not infected with NetNut. This is just Google abusing their monopoly position to hurt its competitors. If apps ship with stealth backdoors to sell access to the user's internal residential network, that's malware. I doubt any users want app providers to sell access to their private file server and anything else on their local network. It doesn't seem like monopoly abuse to exclude such malware from application stores, just like key loggers or apps intercepting other apps network traffic without the user being aware of it (say the banking app's network traffic and password entry).
- charcircuit 3mo ago>sell access to the user's internal residential network That is not what the SDK was doing. The actual code in the SDK protects against this (simplified to take less space): if (addr.isSiteLocalAddress() || addr.isLoopbackAddress()) { LogUtils.e("PopaTunnelAsyncThread", "Hacking? The Host Resolved Ip is " + addr + " on tunnel id:" + tunnelId); throw new IllegalArgumentException("Hacking? The tunnel host resolved ip is internal"); } Local and loopback addresses like 10.0.0.0, 172.16.0.0, 192.168.0.0, and 127.0.0.0 do not work. It will not connect to people's private file servers on their network.
- throwaway7356 3mo agoIt'll still connect to IPv6 addresses and bypass any firewalls. Also users might become part (victim?) of a police investigation because of illegal actions that seem to originate from their local residential connection. So still good to take down such backdoors. Would be nice to go after the botnet operators as well...
- teravor 3mo agoI find the notion that you would use residential proxies to scrape LWN somewhat laughable, I'm reading this article using a VPN. residential proxy bandwidth isn't that cheap, I could see it be used on a reddit (though i would probably just mass register accounts to bypass their block instead).
- rwmj 3mo agoI ran a gitweb server which was battered by bots so I eventually had to take it down. Gitweb! You can just connect using the git protocol and download everything vastly more efficiently! In other words, they don't care at all. For them, residential bandwidth is completely free.
- inigyou 3mo agoRight. This whole controversy makes no sense as a pure scraping thing. It seems more like someone is trying to take the web offline.
- BLKNSLVR 3mo ago> There are ways to tell the difference — the bots usually do not fetch images or CSS, for example — but, by the time that determination is made, the address in question will not be used again. Blocking the address at that point is just a waste of time. Maybe there's no point for the scanned server to block the address, but couldn't collective / shared block lists help with sites that may get scanned by the same address after the initial one? The main problem becomes managing lists of millions of individual addresses. My (only semi-reliable these days, due to lack of time for maintenance) little project has nearly 2.3 million addresses recorded - although only 590k are from 2026, and only 38 were probes on ports 80 and 443. So maybe more manageable than I thought (but my servers don't host anything beyond personal interest to me, and access is filtered via cloudflare, which is it's own "internet control issue"). > In general, these companies range from those that aspire toward some appearance of legitimacy, advertising "GDPR compliance" for example, to others that are just overtly sleazy. Overall, my gut feel on residential proxies is that they're an untrustworthy scourge. I'd be interested in any arguments for residential proxies by people who don't (intend to) profit from using it facilitating them. In regards to Bright Data, one of the companies that attempts to appear legitimate, at minimum these domains should be blocked: brdtnet.com luminatinet.com bright-sdk.com luminati.io As listed in this article, on HN's front page 34 days ago: https://news.ycombinator.com/item?id=48422993 https://news.ycombinator.com/item?id=48422993 (https://blog.includesecurity.com/2026/06/the-smart-tv-in-your-livingroom-is-a-node-in-the-aiscraping-economy/ https://blog.includesecurity.com/2026/06/the-smart-tv-in-you...)
- inigyou 3mo agoWhy would anyone who doesn't have a use for a residential proxy have an argument for residential proxies? I use them to scrape closed sites to make the information more open. For example YouTube.
- andai 3mo agoWell the argument appears to be, people put them in their apps instead of ads. (Or more likely on top of ads.) The argument is money. The users presumably don't know about this, or you know, they clicked, "I agree." Nearly Half of LG Smart TV Apps Contain Residential Proxies https://news.ycombinator.com/item?id=48635954 https://news.ycombinator.com/item?id=48635954
- ArtTimeInvestor 3mo agoIn theory, proof of work that is used to mine a cryptocurrency could be a solution. Bitcoin and others are already secured via massive pow computations. If we could shift that into browsers, no additional energy would be used and we could solve an issue that has been unsolved for too long: How to pay websites that provide useful information other than with ads. The question is which resources typical consumer hardware has that large centralized compute power does not. In-browser POW to pay websites would only be possible if such a resource exists. I am not familiar with the topic, but maybe CPU power and memory? Both seem significant in a typical consumer device. Napkin math: If a consumer device can generate $100 per month, that would be 100/30/24/60/60=$0.00004 per second. If the user waits for 5 seconds before the first pageview, that would then make the website provider $0.0002 per visitor. Serving a million visitors per month is nowadays easily possible on a $10/month machine. So the $0.0002x1000000 = $200 would make the website a nice profit.
- karlgkk 3mo agoProof of work captchas are widely deployed, especially on more niche sites. Kiwiflare is one (used for a harassment forum)
- inigyou 3mo agoYou could've just said Anubis to avoid giving those guys advertisement.
- karlgkk 3mo agoI could’ve said a lot of things. Kiwiflare is the only one I’m aware of (and can remember name of) that provides a visible proof of work occurring in the browser, regardless of reputation scoring (which I don’t think they use). Although, I suppose by definition, anyone using that site is of poor character… If you can think of another example, please be helpful and provide it, thanks.
- Loic 3mo agoProof of work, even "custom", where the user does not need a particular interaction with the page, does not work. The scrapers are running headless Chrome and solving the work. They do not care, they do not pay the bill, the compromised system's owner pays the bill. I have such system for the registration form on one of my website to prevent the double validation of emails to be used to spam emails of victims. The PoW challenge prevents less than 10% of the bots.
- zarzavat 3mo agoThis is a predictable consequence of age verification laws and social media bans. Formerly VPNs were a nice to have but now they are a necessity in many countries to navigate the modern internet. The cheapest way to get a VPN (and if you're a horny and broke teenager perhaps the only way) is to trade your clean but censored IP address for an uncensored IP address in another country. You accept the bot traffic in return, or externalize it to your parents or the owner of the internet connection.
- selfhoster1312 3mo agoThat does not explain why so many residential VPNs operate with so many IPs in countries where there are no social media bans. Here in France: - i know many people who buy shady IPTV boxes from stores/markets for like 50€/year - i know some people who use "smart lightbulbs" and other nonsense - almost everyone i know plays free smartphone games, which as LWN reminded, may contain a shady SDK
- entropyneur 3mo agoSorry, I understand scraping is a problem, but talking about open Internet while simultaneously complaining you can no longer discriminate datacenter IPs like you used to is hypocrisy. I use a datacenter-based IPv6 address because my local ISPs don't offer v6 connectivity and the Internet is already broken for me. And generally the entire idea of a "residential" IP address smells.
- selfhoster1312 3mo agoNoone complained we can't discriminate DC IPs, though to be fair some (imo bad) operators did just that. This is not even about preventing bots, which has perfectly legitimate usecases (eg. Internet Archive). This is about filtering out bad bots/actors who have no respect for your resources and will drain all of it causing bad experience for everyone. But because they know they don't respect robots.txt or even simple rate-limiting, they have to employ so-called residential VPNs. They're residential in that they route through real user connections, and so you can't block the IP/subnet without dropping a certain amount of legitimate human-driven traffic. Personal example: some time ago, i had to disable a wordpress plugin on a site that was causing 100% CPU usage on the whole box (hosting dozens of wordpress instances). That plugin was a simple calendar, but a bot was repeatedly scraping non-existent (or rather, "no event planned for this day") pages for every date in the calendar that you can represent in the DB timestamp, clearing the cache as it went to try and find new events for 1000 years ago. Whoever operates this IP space doesn't matter to me, i'd just like to block them because they don't respect robot.txt… but i can't because they use a "residential proxy" and will change IP address every hour or so.
- inigyou 3mo agoYou could always try various counterattacks - returning a repeating compressed stream that decompresses to several terabytes, or returning one byte per second, or returning an endless list of hyperlinks to a honeypot, if the date is older than 200 years.
- selfhoster1312 3mo ago
- klamann 3mo agoI think this Anubis project is a terrible solution to the problem posed by aggressive web scrapers. Using a web browser with reasonable privacy settings has become a big loss in quality of life already, but the first time I encountered Anubis I got completely locked out of most web servers that deployed it. The situation has improved a little, but I hate that maintainers of great web services have rationalized themselves into believing that creating massive barriers to access their sites is a fair trade-off. Unsurprisingly, I have nothing but negative associations with their mascot. The FSF has the right idea about all this: > Some web developers have started integrating a program called Anubis to decrease the amount of requests that automated systems send and therefore help the website avoid being DDoSed. The problem is that Anubis makes the website send out a free JavaScript program that acts like malware. A website using Anubis will respond to a request for a webpage with a free JavaScript program and not the page that was requested. If you run the JavaScript program sent through Anubis, it will do some useless computations on random numbers and keep one CPU entirely busy. It could take less than a second or over a minute. When it is done, it sends the computation results back to the website. The website will verify that the useless computation was done by looking at the results and only then give access to the originally requested page. > At the FSF, we do not support this scheme because it conflicts with the principles of software freedom. The Anubis JavaScript program's calculations are the same kind of calculations done by crypto-currency mining programs. A program which does calculations that a user does not want done is a form of malware. Proprietary software is often malware, and people often run it not because they want to, but because they have been pressured into it. If we made our website use Anubis, we would be pressuring users into running malware. Even though it is free software, it is part of a scheme that is far too similar to proprietary software to be acceptable. We want users to control their own computing and to have autonomy, independence, and freedom. https://www.fsf.org/blogs/sysadmin/our-small-team-vs-millions-of-bots https://www.fsf.org/blogs/sysadmin/our-small-team-vs-million...
- InsideOutSanta 3mo agoI think all of these mitigations are unfortunate. They hurt one of the things that makes the web cool: it's a stable, stateless, idempotent way to access data. This makes it a prime target for aggressive scraping by LLM companies, but it also makes it accessible and fast, and a prime target for benign use (like archive.org or "read later" services). For my own sites, I'll eat the cost of the crawlers (mitigated by making the sites as efficient as possible) and keep them available to everyone.
- m00dy 3mo agoDisclaimer: I own and operate proxybase.xyz [0] Hi HN, I wanted to jump in and share a few thoughts. Not all the residential networks mentioned on this page are bad actors. Transparency & Auditing: Our clients are completely open-source, and we run strict internal audits before every single release [1][2]. You don't need to be a security wizard to verify this, either. You can easily audit the code yourself—just clone the repo, feed it into an AI, and ask the right questions. Ethical Sourcing: Consent is everything. At Proxybase, we always get explicit consent from our providers before adding them to the pool. This is exactly how ethical sourcing should be done. Historically, this industry has been incredibly shady think malware bundled into iOS/Android apps or second-tier smart TVs secretly installing background scrapers. Fortunately, the sector is finally becoming more ethically aware. Fair Payouts: A lot of networks hold onto provider funds for months, staking them to earn passive income while making users wait. Between sky-high payout thresholds and endless waiting periods, it’s a broken system. At Proxybase, we have a $1 minimum payout sent directly to your wallet using US stablecoins. If you have a better idea on how we can make this industry better, just lmk. I'm reading/writing on HN everyday. [0] https://proxybase.xyz https://proxybase.xyz [1] https://github.com/proxybasehq/proxybase-gui https://github.com/proxybasehq/proxybase-gui [2] https://github.com/proxybasehq/proxybase-cli https://github.com/proxybasehq/proxybase-cli
- mrenzo 3mo ago[flagged]
- nikolife2016 3mo agoRunning a small public JSON API, the traffic breakdown is eye-opening. Roughly half is trust/uptime "scanners" and generic monitors; a solid chunk is well-behaved crawlers that declare themselves with real UAs and honor robots; and then there's a long tail of vuln-scanners blindly probing for /.env, /.git, wp-login and the like. The genuinely evasive residential-proxy scraping is a minority by volume but by far the hardest to separate from real users — it's the one bucket where UA and IP both look residential, so you can't tell bot from human without behavioral signals. What's shifted in the last year: the "polite" bots got politer, while the abusive layer moved almost entirely onto residential proxies. IP reputation alone is basically dead as a filter now.
- CodesInChaos 3mo agoHow much does routing traffic though residential proxies cost?
- inigyou 3mo agoYou can just Google this. It isn't illegal, you can find many providers, you can even pay with your credit card. Usually around $0.20/GB - other types of proxies are cheaper.
- jappgar 3mo agoI was involved in both sides of this battle over ten years ago. Things haven't changed all that much. It's important to note that neither side has moral legitimacy. Not everyone who carries a rifle is a enemy. Not everyone wearing body armor is a saint. I have given up on the idea that "human vs bot" matters at all when it comes to anything other than voting (which should only be done in person with paper and pen, by the way.) You could make an argument that "likes" are a form of voting, but you shouldn't. We need to abandon the idea of supposedly democratized algorithms and focus instead on actual democracy.
- georgyo 3mo agoThe article at the end talks about how is very easy for arbitrary apps from app stores can install a residential proxy on your phone. 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. And both ios and android make it hard to deny apps network access. But really, this finally explains the hordes of really basic boring games that just advertise other boring games. Idle games and the like that really just want you to keep your phone unlocked and open. Millions of downloads on the app stores for entirely offline content (and ads) and no way to block the network access.
- dannyfritz07 3mo agoGrapheneOS allows you to deny network access per app pretty trivially. Google Play services make it a bit more difficult because the app might marshall the network request through that; I'm not sure how to verify that behavior when it happens.
- igoose1 3mo agoThanks for mentioning GrapheneOS. I'll just explain to others what "pretty trivially" means here. GrapheneOS adds a huge checkmark "Network access?" when you install an app. It's impossible to miss.
- bombcar 3mo agoThe problem is the yes/no options - I want to give access to specific endpoints - but now that’s “all of Cloudflare for everything” which means the web entire. It used to be you could let Onkyo App™ access the five IPs for onkyo.com and be done with it.
- inigyou 3mo agoI think Google Play Services will only marshal data to and from Google. This will bypass the permission if you're using it as a firewall, but won't let the app run a residential proxy.
- 3mo ago
- phendrenad2 3mo agoEver since bots became a problem on the internet 10-20 years ago, it has seemed like the common-sense solution is some kind of micropayment. Pay $0.01 to view the page. When money is on the line, scrapers are likely to be more well-behaved, even if they do pay. The problem is, and has always been, the friction of payment. How do you pay $0.01? The credit card processors will tack on a $6 surcharge. We need a trusted third-party that turn money into "internet article credits" that you can spend in small increments, like a video game. But I suspect that thousands of people have already though of this system, and tried it, but ran into some roadblock. I'm guessing there's some egregious regulation that makes micropayments impossible.
- RetroTechie 3mo ago> I'm guessing there's some egregious regulation that makes micropayments impossible. More likely there isn't any kind of universal standard that's easy to implement for browser makers, has low overhead, and preserves internet users' anonymity as much as possible. The currently existing friction of using micropayments is the problem here, I suspect.
- inigyou 3mo agoProbably with something similar to Lightning Network, which reallocates pre-committed funds between two or more parties. But not with Lightning Network itself, for several reasons including how costly it is to pre-commit the funds.
- setheron 3mo agohttps://fzakaria.com/2026/07/09/who-does-anubis-actually-stop https://fzakaria.com/2026/07/09/who-does-anubis-actually-sto... I wrote about this recently as well.
- dlenski 3mo agoAs this article points out, it's tremendously unclear who is using residential proxies. The big AI models claim they're not using them. I'm not inclined to "just believe them", but no incriminating evidence has leaked, and—as pointed out in the article—many of the bots that are running on these residential proxy botnets are coded in incredibly stupid and inefficient ways. How confident are people who research this stuff that the RP botnets are actually being used for AI training?
- xena 3mo agoThat's the only theory we have that doesn't sound like a conspiracy theory. The only other credible ideas are that someone's doing some kind of dataset arbitrage by scraping the fuck out of everything and selling companies data that is technically new (by means of the scrape date being newer).
- inigyou 3mo agoOr someone is trying to DdoS the entire web.
- dlenski 3mo agoRight, and as I understand it the timing also lines up: these ill-behaved scraper-bot-nets exploded along with GenAI in the last 3-4 years. Still, it seems to me like something doesn't add up. Running these botnets is perhaps cheap but it isn't free, and dumping all their data into LLM training is truly expensive: would so many of these bots be so blitheringly inefficient in their scraping patterns if all of the results were getting fed into LLM training? Have there been no leaks or whistleblowers from the "semi-legit" RP brokers?
- 1vuio0pswjnm7 3mo ago"The startup, which markets itself as the second-largest data collection firm after Alphabet Inc.s Google, has grown significantly on demand from data-hungry AI companies." https://www.bloomberg.com/news/articles/2026-07-10/web-scraper-sets-1-million-bug-bounty-as-industry-scrutinized https://www.bloomberg.com/news/articles/2026-07-10/web-scrap... Seems like Google is going after the competition
- nektro 3mo agoi wonder if residential ISPs can play a bigger role here
- subarctic 3mo agoI don't run one of these sites that has these issues so I'm really not aware of this problem. How can it be that sites are getting overwhelmed with scrapers that are just looking for training data? You only need to scrape it once to train a model, so shouldn't there be less traffic from this than there is from search engines? On the other hand if the article is wrong and the traffic is coming from other ai uses (like an agent visiting pages on behalf of a user) then that would make sense.
- tancop 3mo agoi think the best way to keep your site working for legit users is serving static cached pages to suspected bots. decide based off something like cloudflare bot score. crawlers get redirected to content that might be stale but its still useful for them and costs you almost nothing to serve. put a warning banner on it so if a user (or smart agent) accidentally ends up on the bot version they can click on it and get anubis checked for the real site. of course this only works for people like me who want their work to be used for training ai.