9 ms·
The relevance of IP addresses in the tracking ecosystem [pdf]
- jentulman 6y agofrom the abstract.... "In this paper, we study the stability of the public IP addresses a user device uses to communicate with our server. Over time, a same device communicates with our server using a set of distinct IP addresses, but we find that devices reuse some of their previous IP addresses for long periods of time. We call this IP address retention and, the duration for which an IP address is retained by a device, is named the IP address retention period. We present an analysis of 34,488 unique public IP addresses collected from 2,230 users over a period of 111 days and we show that IP addresses remain a prime vector for online tracking. 87 % of participants retain at least one IP address for more than a month and 45 % of ISPs in our dataset allow keeping the same IP address for more than 30 days."
- annoyingnoob 6y agoI suppose this only works if you combine IP with some other information, like username or browser fingerprint. Otherwise you could be tracking multiple 'users' at the same IP.
- parhamn 6y agoThe whole multiple-users-on-one-ipv4 thing always feels like it distracts from what a tracker's goal really is. It is definitely beyond sufficient identification and more than needed for a tracker to start targeting you with ads and what not. In some ways the fact that IP only tracking is lossy but still good enough is something to be afraid of. It is easy to 'taint' an IP and your households behind-the-scenes profile over a google search by your guest using your wifi.
- annoyingnoob 6y agoAdvertising that is relevant to my wife probably isn't relevant to me, using IP is not good enough. The title and focus here is on user tracking, which is the goal I'm commenting about. Its quite common to have multiple users behind the same IP, those users and their actions may or may not be related (I have no idea what my co-workers might be doing for example but we share a static IP). I honestly didn't follow your Google search/wifi example.
- pm_me_ur_fullz 6y ago"It's listening to me!" and your spouse's typed in search queries and your playstation and your smart coffee machine and alexa actually is thats all without getting me started on the analytics networks in other apps sharing with the same data brokers between apps
- virgilp 6y ago> using IP is not good enough Define "not good enough"(for who is it not good enough? for advertisers?). I used to work on a project linking cookies together in anonymous profiles (for advertising), initially the plan was to separate household/individual profiles using different heuristics, but I think eventually it turned out that nobody cares that much - the advertisers just wanted a rough cross-device profile, not perfect accuracy. I mean sure, for marketing purposes (as well as engineering/ "performance" reasons) you needed figures to brag about accuracy and whatnot. But it's really really hard to get those figures right, and ultimately the thing that will convince advertisers is "using cross-device profile improved conversion rate by 10%"; everything else is a detail in comparison.
- annoyingnoob 6y agoThe title and focus here is on user tracking, which is the goal I'm commenting about.
- virgilp 6y agoI have no idea what you mean by user tracking, can you explain?
- annoyingnoob 6y agoA user is usually defined as a human being. I'm a human being and I did not come with a network connection. The world, generally, does not assign IP addresses to human beings - thus an IP address cannot track me as a human, a human computer user. The paper/study here used a combination of IP address and a UUID to track 'users' which are presumably but not necessarily individual humans. It was the use of this UUID that allowed the authors to draw conclusions about how often a given user is seen at the same IP address. The IP address by itself did not provide enough tracking to draw the conclusions in the paper. Is IP address identification 'good enough' for advertising, maybe. But IP is no where near definitive for identifying individual humans. The paper does not say how often they found multiple 'users' using the same IP address. The paper focuses on individual reuse of IPs. How often did they find different users using the same IP in their study? In internet advertising no one expects every ad to trigger a sale. Since there is some amount of acceptable waste in the system this paper has decided that IP is 'good enough' for targeting - because its okay if we are throwing away 20% or more of the ads (we didn't expect any traction from them anyway seems to be the thinking). And the authors admit that a larger or different dataset might show different results.
- kube-system 6y agoThe big conclusion to be drawn from this paper is that IP address tracking can be combined with client-side tracking techniques to perform reidentification.
- Zenst 6y agoMost users will at the very least use two IP addresses - home broadband and mobile SIM broadband. Then you have wifi hotspots, friends wifi. The average user uses many IP's and not limited to the range of one ISP. SO whilst you can fingerprint devices and usage patterns, the IP address will by itself be useless to identify such users, it may well augment a little but is no solution. But then IPv4 shares many IP addresses across mobile and broadband users in various ways. Most do not have a fixed IP and even those that do, do not have a fixed IP upon their mobile data activities - unless they VPN into their home broadband. Though if they use service VPN offerings, then another layer of IP ranges. So the potential towards false assertions based upon an IP and user usage may well trip up and fail. Imagine using an ISP with dynamic IP and the next user of that IP uses it for crime, well with some bad logging and aggressive association, you can mislabel somebody for a crime they did not commit. Roll on IPv6 and with that, mobile carriers would of been the obvious benefit of that, yet I'm not aware of any progressing that in any timely manner and chug along using a pool of IPv4 and various tricks to make those cater for many.
- scared2 6y agoHowever in this paper the authors tried to show is it's stability over time. So overall their "findings" indicate ip addresses should not be overlooked in privacy protection.s they stated as follows: "... Over time, a same device communicates with our server using a set of distinct IP addresses, but we find that devices reuse some of their previous IP addresses for long periods of time. We call this IP address retention."
- Zenst 6y agoAs somebody who tracks and logs their IP's given via ISP etc, I can attest, not that distinct over time.
- jedberg 6y agoIPv6 improves this situation (now). At first, ipv6 was actually a lot worse, since the back 1/2 of your address was your MAC address, allowing your device to be tracked around the internet no matter where it went. People quickly realized this flaw, and updated the standard so that basically your client gets to pick the second 1/2 of your address now. And the nice thing is, most major platforms will actually run multiple addresses in parallel, allowing new connections to use a new address while old connections keep using the old one. So while ipv4 adds some protection by having multiple clients behind the firewall, ipv6 actually makes it better by looking like even more clients behind the firewall. Combined with a browser that blocks fingerprinting you get slightly better privacy with ipv6.
- saurik 6y agoThat seems backwards: with NAT you couldn't identify all of the individual computers on my Internet connection; but now, with IPv6, you either can at worst (as every device has its own IP address that it reuses) or, at best (generating a new address for every single connection), are just getting yourself back to where you were with NAT. I appreciate that for a while IPv6 was actively much worse as it allowed address correlation across multiple networks (due to the MAC address being the same in the lower 64 bits), but fixing that doesn't make it better than NAT: if you mostly care about privacy, it still seems to make the most sense to use NAT if at all possible.
- ignoramous 6y ago> with NAT you couldn't identify all of the individual computers on my Internet connection... Not if you're using WebRTC which would promptly leak the private IP address: https://news.ycombinator.com/item?id=12528184 https://news.ycombinator.com/item?id=12528184
- zrm 6y agoThere are also several other methods that often work, e.g. the X-Forwarded-For HTTP header if there is a local proxy, or a long list of non-HTTP protocols or tunnel-XYZ-inside-HTTP protocols that have the client IP as part of the protocol. Temporary IPv6 addresses actually solve this, especially for P2P systems that do benefit from knowing the client IP in case there is a peer on the same LAN, because the address they get is the same one the remote server sees (no additional information) and then it's both a valid local address but hard to correlate with anything when every device can have hundreds of them at once.
- AndyMcConachie 6y agoI worked on a document a few years back on anonymizing IP addresses. If you find yourself in a situation where you need to balance anonymization of IP addresses with research needs, this paper may be useful. https://www.icann.org/en/system/files/files/rssac-040-07aug18-en.pdf https://www.icann.org/en/system/files/files/rssac-040-07aug1...
- api 6y agoIP is just one more data point. There are already so many ways a browser can be fingerprinted, it doesn't make things that much worse. While you can limit your exposure a bit, I long ago reached the conclusion that strong privacy is impossible in the current client/server web model. There is too much surface area.