3 ms·
> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really. So if Google has distilled someone's emails ov
by anf 9y ago
> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really.
So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?
- spydum 9y agoAs long as they retain no data which could specifically identify the original person, yes. There is nothing wrong with building segmentation models as long as they aren't specific enough to identify a specific person. My concern would be, how granular is too granular? What if we added "and live in zip code 12355 and is registered Green Party"? This now gets eerily specific, and might be sufficient to identify an individual.
- anf 9y agoWhy would they ever discard that? Why would there be a granularity where ML suddenly stops working? Why would you even stop at one model per person, instead of one model per mood, or modes of thought at different stress points?
- creaghpatr 9y agoIn fact they would desire that granularity most of all so as to reconcile the past and future state psychographic profiles for an individual- then they could attempt to isolate the causation of a state change- basically they need to identify the moment an individuals profile reflects the change from democrat to republican or vice versa. Or Religious to atheist etc.
- darawk 9y agoIf they kept information like that, then yes that would be an invasion of privacy. But that sort of information is almost certainly not encoded in an ML model trained on 50 million people's data.
- anf 9y agoHow do you know what CA trained on, or what's possible? Do you have qualifications in ML?
- darawk 9y agoI know what they trained on because it's been reported on. They got around 50 million people's FB profiles, and a smaller subset's (300k, I think) personality test results. I use ML models every day in my work, and understand how they function. It is true that individuals information is probabilistically encoded into the parameters of the model. However, if the model is any good, the people they trained on's information is encoded only a bit more than that of the entire population. There is sort of a privacy issue in the following sense: The models they've built have learned relationships between preferences and personalities that they wouldn't otherwise have been able to learn. But these relationships are abstract. They are not tethered to any particular, identifiable individual. A reasonable argument can be made that those learned relationships are, in a sense, stolen property. And I think arguments along those lines are interesting things that we'll have to explore as this sort of thing becomes more common. But the idea that this model invades individuals privacy just isn't really true.
- pflanze 9y agoBut if the resulting model doesn't contain information about individuals, how does this help targeting individuals for the campaign? Edit: is it that the model is then applied to only strictly public data about the person? If so I guess the interesting question then becomes whether the model is definitely not anything near overfitting (i.e. containing enough information to match a person's public data directly since it was trained on it (amongst other data))? (I'm not an ML developer.) Edit 2: also, going with your comparison with the "20 most representative pixels", it seems interesting then that 'this much' (although not exactly sure how much) information can be inferred from a public profile when just also knowing enough about the whole Facebook population. OK, so perhaps a human would be able to infer about as much, but doesn't scale, and that's why the model becomes valuable?
- jcranberry 9y agoThat's a ridiculous response. If they managed to infer this characteristic from emails, what they would keep is a tool which, given that set of emails again, infer the same characteristics (and theoretically a similar set of emails). They would by no means be allowed to keep the kind of information you described. What is more relevant is a model which, given characteristics such as "closeted homosexual with a deeply repressed leather fetish", they would be able to infer other characteristics, such as support of particular political candidates, responsiveness towards targeted political or commercial ad campaigns, etc. That's what's relevant here.
- caseysoftware 9y agoSince the source data was deleted - according to current standards and policies - their hands are probably technically clean. But there may be another angle of attack. In the US, you're not allowed to benefit directly from a crime you committed. For example, if you rob a bank, you can't buy your mother a car with the money and say "sorry, it's gone!" when the police come knocking. With that line of reasoning and if there was a legal, privacy, or at least a TOS breach in collecting the data, the derivative machine learning models may be tainted also. Then again, it's likely impossible to prove exactly what data went into the model, so hard to establish which models might be tainted.