3 ms·
An interesting way to look at this, but wouldn't it be possible to pluck this pattern out from relevant data?
by iask 10y ago
An interesting way to look at this, but wouldn't it be possible to pluck this pattern out from relevant data?
- danbruc 10y agoIt most likely is because it is incredible hard to create realistic fake traffic. The moment you start using such a tool is probably very easy to detect due to a sudden jump in traffic. Then you can use prior data to figure out what you were actually doing - sites you visit, times you are active times, how you switch between sites, how long you stay on sites, and so on. Randomly following links will show very different distributions. You have to fake your browsing behavior with pretty high accuracy but if you do that, you defeat the entire purpose, you just created a copy of your browsing behavior. Maybe if there is no prior knowledge of your browsing behavior and the fake traffic is human-like enough, but even in that case it seems not too unlikely that one could separate the two traffic sources. If you want to hide your browsing behavior, the fake traffic must be different from yours. But if there is a difference, the difference can potentially be exploited to separate the sources.
- state_less 10y agoI was thinking the same thing, real sessions might be separable from fake. This smells like an adversarial problem, where one AI makes fake traffic and one tries to detect it. Perhaps the fake session creation bot could mingle on the same sites as the real sessions to make it harder to distinguish.
- jerf 10y agoIf I were writing this, I would try to get it into as many hands as possible, and do some sort of backfeeding of data to generate "real" profiles of user activities, then send back down suitably randomized profiles of other real users that you use for some period of time. It's important to conceive of this as an arms race, and I think one of the underappreciated things about arms races is that if you successfully anticipate the next three things your opponent will do and counter them, you can end up discouraging them from even fighting in the arms race. I don't know how random the page visits are, but if they are random in, well, almost any sense, eventually they will be something the trackers can filter out. "Well, this person shows a really strong signal for football, and weak signals for anime, robotics, accounting, and a hundred other things. Show them the football ads." Even today that wouldn't necessarily pose much of a challenge. By contrast, if you get profiles from volunteers that are real, and then start mixing them up so that everybody downloading this extension shows ten or fifty equally strong signals of interest, of which only one is real, that'll scramble the data being collected something fierce, and require the data collectors to jump straight to very sophisticated teasing out of what's really true, which initially won't even be worth it until a lot more people are using this stuff. There's a lot of elaborations you can think of from there (temporal correlations, i.e., college basketball interest should be spiking now, vs. football), etc. I have no idea what this extension is doing, because I can't seem to find any data about what it is doing, so maybe it's doing some of this stuff, but I expect they'd be talking about at least the data they'll need to collect to make this really work if it was going to work this way, so I assume without proof that it is not this. In general, it's worth pointing out that a lot of systems can't be fooled by uniformly random data, because all real-world systems already have to be able to filter that out because all real-world systems experience noise, of which at least a significant component is probably more-or-less uniformly random. If you really want to scramble a system you need to be more clever.
- danbruc 10y agoBy contrast, if you get profiles from volunteers that are real, and then start mixing them up so that everybody downloading this extension shows ten or fifty equally strong signals of interest, of which only one is real, that'll scramble the data being collected something fierce [...] This would also increase your traffic ten or fifty fold in a first approximation. You would make a single household look like a family with several people sharing the Internet connection. Netflix for example can separate several people using the same account, I read about that during the Netflix Prize [1]. But overall I agree with your point, you probably can throw a wrench into the machinery maybe even to the point that it is not worth to pursue tracking any longer. But it certainly is not easy and randomly clicking on links will definitely not do the trick. [1] https://en.wikipedia.org/wiki/Netflix_Prize https://en.wikipedia.org/wiki/Netflix_Prize
- jerf 10y ago"This would also increase your traffic ten or fifty fold in a first approximation." I suspect if you worked at it you could mathematically prove that will be a necessary condition for any effective data collection scrambling technique, under an assumption that we can't use proxying to other nodes. And I do mean "mathematically" fully literally. I'm eliminating the possibility of proxy, where you try to set up a situation where you create a P2P network and trade page views around, because I think that only works with a really abstract view of how the tracking works. In practice, as soon as you want to use a site in an authenticated fashion you're getting tracked via that authenticated account, so I think I can argue the only real possibility for scrambling the data is for it to source from the same network location as the real data you are generating. On that note, it occurs to me this plugin probably ought to be automatically creating authenticated accounts on services you don't care about; the authenticated status of an account creates a shining signal that may be too bright to mask. Considering that a lot of the data you're trying to mask would be coming from Facebook, that would be a problem. And now that I think of that, use of this plugin really ought to result in your Facebook profile getting pretty badly scrambled, too. Man, this is a big challenge. There's a part of me that is actually quite sad I can't drop everything and start going to work on this right now for pay. This sounds wicked fun, but way bigger than I could ever dream to take on.