5 ms·
It seems routine to see a bunch of browser User-Agents from the same IP
- xena 2y agoSome of the Mastodon traffic is because of https://masto.host https://masto.host
- burningburned 2y ago[flagged]
- hackernewds 2y agowhat's nat?
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- thayne 2y ago"network address translation"[1] Basically, multiple devices access the Internet through a gateway that uses a single ip address, using different ports to map back to the local ip address. [1]: https://en.m.wikipedia.org/wiki/Network_address_translation https://en.m.wikipedia.org/wiki/Network_address_translation
- p_l 2y agoAnd if your web server is on v4 only, more and more phones will access it through NAT64/DNS64 from a shared pool of addresses
- russianGuy83829 2y agosorry about the off topic, but has anything come out of this? https://news.ycombinator.com/item?id=39142069 https://news.ycombinator.com/item?id=39142069
- p_l 2y agoUnfortunately not yet. I had more pending (paying) needs that conflicted with the patches necessary to support the AMD driver, and I didn't have time to backport the necessary bits to 6.7. This is no longer a problem, so time permitting I might try something in this area next month - need to deliver a freelancer project first ;)
- russianGuy83829 2y agogood luck to you on your endeavors! I’m sure the HN crowd will be interested in whatever comes out of this ツ
- deleted 2y ago[deleted]
- gnfargbl 2y agoIP addresses vary in scale in terms of the number of devices behind them. There are many, many reasons for this including corporate gateways, mobile networks, CGNAT, VPNs, public wifi and yes, even home NATs. There's no "normal" UA cardinality count that you can use for filtering. However, there may be other traffic characteristics that are useful; the shape of the distribution, for instance.
- dz0ny 2y agoCGNAT :)
- af78 2y ago= https://en.wikipedia.org/wiki/Carrier-grade_NAT https://en.wikipedia.org/wiki/Carrier-grade_NAT
- whalesalad 2y agoCrawlers.
- nutrie 2y agoCrawlers? I'm behind a fixed IP, along with thousands of other people who happen to be customers of the same ISP.
- dylan604 2y agoTo anyone that has ever tried to parse their http server's access logs, this is something that they realize and come to the same conclusion. This is one of the reasons why third party statistics like GA exist. If this was a blog from someone that this was their very first attempt ever at parsing logs, this might seem more reasonable. However, this person claims they know parsing access logs is frustrating and yet claim this was a novel/clever idea which does not compute. Just a simple web search on issues parsing access logs would show this is not a new idea nor would it be a successful one.
- tootie 2y agoThis is a big issue with podcast analytics since most downloads come via podcast apps that don't allow any client code at all. We only have server logs to go off of and hence problems like NAT are essentially unsolvable.
- throw1902378190 2y ago> One of my recent clever ideas was to look for IP addresses that requested content here using several different HTTP 'User-Agent' values I'm always surprised when a blog post like this makes it to the front page of HN. I suspect that their only experience with web hosting is this blog as there's nothing clever or interesting about their experiment.
- dylan604 2y agoI always chalk it up to those that refuse to look at history are doomed to repeat it.
- donatj 2y agoWe sell to K12 schools. Depending on the district sometimes they'll have a single IP for an entire district. If not, often a single IP for the building. This makes things like DoS protection... Interesting as we can regularly have thousands of users legitimately using our stuff from the same IP concurrently.
- thayne 2y agoSame thing applies to large businesses where a large number of employees share a single (or maybe a handful of) ip address for a corporate network.
- deleted 2y ago[deleted]
- jonatron 2y agoTLS fingerprinting can spot things like requests made with python requests easily.
- klabb3 2y agoIf the goal is to have a decentralized web without the Cloudflare protection racket, would proof of work be feasible and/or desirable to deter mass crawling? At the heart of the issue is that sending a http get / is cheaper than generating the response, tilting the economics in favor of the content hoarders. Or is mass crawling of static content completely fine and we should just be better at caching?
- tatersolid 2y ago> would proof of work be feasible and/or desirable to deter mass crawling? No. Bad actors generally don’t pay for CPU/GPU as they use compromised hosts or legit hosts with stolen payment info. This same “proof of work” solution didn’t fix email spam 20 years ago, for the same reasons. It only adds costs to legitimate activities while being a minor inconvenience to bad actors.
- eurleif 2y agoCompromising a host, or stealing payment info, cost money, or at least time. Presumably they cost less than you get, but it's not free. If you can increase the CPU time it takes to send a spam message by 10x, then the cost to send a spam message will increase by 10x. This should be especially valuable for as long as there's a "not outrunning the bear, just outrunning your friend" effect, but even with full adoption, I would expect to see spam meaningfully reduced. Was this solution actually shown not to work 20 years because spammers are able to pay the cost? My sense is that Hashcash never saw widespread adoption, and we don't really know how it would have worked out.
- jcynix 2y agoAs you wrote, Hashcash never saw widespread adoption, sadly. The big players all came up with their half-assed "alternate solutions" and what I saw was that spammers where the first to adopt DKIM etc. I still use greylisting which keeps a remarkable number of spammers away, because they do not invest the resources to wait for n minutes for a retry. And I refuse email from clients without a proper reverse DNS entry, which more often than not appear to be hacked machines, or certain "organisations" in countries far away ...
- elpocko 2y agoMy browser is configured to send fake referrer and user agent strings. I tried not sending a UA header at all but it broke too many websites.
- renegat0x0 2y agoThe Internet is full of bots. It is full of corporate bots that scrape data for AI. It is full of government surveillance. It is full of North Korean Lazarus hackers. It is also full of tech enthusiasts that have various projects. It is full of people that use various RSS readers. I often am repulsed that servers reads user agent. It is often not used for compatibility reasons, but just to detect if you are running official browser, or if you are 'privileged user'. This leads, obviously, that some programs needs to use disguise, which would not happen if servers were friendly toward friendly bots. Imho there is a problem with bad-bots, not all bots, that take away your bandwidth, processing power, etc. etc.
- ollybee 2y ago"I often am repulsed that servers reads user agent." That is a backwards way of looking at it, they dont take the user agent, you give it. You you give any user agent you like or none. No one providing a free resource online has any obligation to treat your requests in a particular way. If you dont like the way a particular HTTP service is responding, then dont send it requests.
- Joker_vD 2y agoThe thing is, I've seen quite several web servers that refuse to serve your request unless you provide any User-Agent, and I mean any: it can even be "}__test|O:21:\" (which is not a valid UA), but it can't be missing. Yeah, sure, very security, much protection.
- Thoreandan 2y agoHas anyone had success using things like the rsync or bgp feeds from Spamhaus to drop this traffic?
- throwaway2016a 2y agoWhile I suspect this may be old news to a lot of the HN crowd, this article interesting information that a lot of people may not know. I ran a website years ago that was targeted towards college students and would pretty consistently see many (hundreds) different UAs under the same IPv4 address due to NAT. Let alone proxies, VPNs, etc. Yet every once and a while someone will suggest the idea of rate limiting based on IP so it's definitely not universal knowledge even among developers how common it is for multiple users to share the same IPv4.