3 ms·
I think you might have missed the deidentification piece?
by jefftk 2mo ago
I think you might have missed the deidentification piece?
- reaperducer 2mo agoNo such animal.
- fileeditview 2mo agoYou have to trust that this really "deidentifies". Time and time again it was shown, that the measures taken were not enough to anonymize. E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?
- piva00 2mo agoEven if they follow to the letter a deidentification process, Google and Meta have so much data about individuals that re-identification shouldn't be very hard for the majority of airline passengers' data they put their hands on. Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
- xp84 2mo agoThe idea that there’s a nefarious plot to do something super evil with this data is a bit crackpot though based on their incentives. Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.” If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.
- piva00 2mo agoNo need for a nefarious plot, as usual with capitalism it just needs incentives. They will have the data, if at some point it's beneficial to Google to re-identify it then it will be done. Do I believe they have incentives to do it now? No, as you point it out for their advertisement cash-cow they can already just rely on their own data (GMail, Search, Google Flights) but nothing stops them from the potential later on, the data is now theirs. Likely its value is just to train LLMs but the funny thing about data is that you can always try to find ways to extract more value out of it. I'd prefer there was no possibility for that without requiring me to trust Google (or any corporation). In an ideal world my data would be mine to control, not to be traded in deals among 3rd parties, it's valuable and I've spent time generating it so in a sense I've done free work to be extracted by these corpos.
- xrd 2mo agoNot trying to be snarky, and perhaps it wasn't well stated, but the last paragraph I said I'm concerned about identification via my writing style. If they have my emails, they would have my writing style. It doesn't have to be tied to PII there, they can cross reference it with my blog. I'm speculating because I read that you can identify people by a few sentences of their writing. "Deidentification" seems really murky and imprecise at best.
- jefftk 2mo agoReidentification via writing style is definitely possible, and I doubt the vendor will modify things in a way sufficient to handle that. But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
- maybewhenthesun 2mo agoBased on their trackrecord, That's definitely a concern. I don't really understand on which basis you conclude 'isn't one of them' . 'Don't be evil' ? :-P
- jefftk 2mo agoI think this is the kind of place where applying bounded distrust is critical: it's not whether we trust Google overall, it's about figuring out what sorts of statements we should expect to effectively bind companies and in what ways. For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.) On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it. If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data. The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.
- Leonard_of_Q 2mo agoI have a bridge for sale, hardly seen use, pay me ${money} and you can collect it in New York City. Interested?
- Larrikin 2mo agoEven before LLMs there were multiple papers written about ways to to reidentify people with ML and other statistical analysis. It is probably now even more trivial especially if you are Google.