4 ms·
With that information I can: • Say with a lot of confidence that billy69420@cbt.eu is owned by one person, called William. • Infer that, since the e-mail says B
by dafoex 5y ago
With that information I can:
• Say with a lot of confidence that billy69420@cbt.eu is owned by one person, called William.
• Infer that, since the e-mail says Billy, but the name says William, the latter is probably what appears on any official documentation.
• Infer that any services signed up to using billy69420@cbt.eu as the e-mail address are more than likely tied to this William person - including any services I run or otherwise have administrative access to.
• Search for billy69420@cbt.eu on various public information repositories and reverse email searches.
• Add an entry to a database using billy69420@cbt.eu as a primary key (e-mail addresses are necessarily unique) to enable me to keep track of my current and future information and inferences, thus building a profile on the person.
I can also understand that the sex and weed references mean that you've probably not given a real e-mail in this example, but if it were real I might be able to infer an age range (or at least a level of maturity) of the user from that, although this would be unreliable as the user could have made the e-mail account some time ago and just still happen to be using it.
- bserge 5y agoFrom OPs link to the official guidelines: "If data are inaccurate to the point that no individual can be identified, then the information is not personal data. (e.g. If you refer to “the man who lives at 12 Mulberry Lane had a party last night,” when Mulberry Lane ends at number 10, that’s not personal data.). I could guess "12" was a mistake and just send spam or visit the man living at number 10. That's way more dangerous imo than an old skool email address and a first name. Yet it's apparently not considered personal data. That's why I asked, actually, since it confused me.
- currency 5y agoPlenty of addresses are designed to mislead. Twitter bots, especially. You can't actually assume good faith on the part of the owner of the address.
- TeMPOraL 5y agoYou can at scale. Separating bots from real people is a somewhat orthogonal concern, and given a large dataset containing mostly real people, you can safely assume vast majority of the data is not purposefully misleading. Vast majority is good enough for advertising, and mistakes in non-advertising use usually create a problem for the e-mail owner, not the database user.