3 ms·
I'm curious how you can know the tool is 95% accurate, if it's being tested on real world data, such as from reddit etc.? I can only assume it was tested on a
by deutronium 10y ago
I'm curious how you can know the tool is 95% accurate, if it's being tested on real world data, such as from reddit etc.?
I can only assume it was tested on a synthetic dataset perhaps.
Also I'm wondering how many unique users are present in the dataset, along with the volume of content for each user.
- gwern 10y agoYou can easily turn any dataset with labeled authors into a de-anonymization dataset: split each author's writings in half and give them different IDs. Now you know the true answer for every pairwise combination.
- deutronium 10y agoYeah, that's a good point. So with that approach you could even get results of accuracy from data from a single social media source at least I guess.
- firebones 10y agoThe problem with that approach is that it is training the classifier to solve a different problem than what is claimed. It assumes that there is no difference between how people write when they are identifiable and when they are not (since it would only train on identifiable samples that are anonymized after the fact). Further, it would require solving the problem of identifying authors between different media--which would be a huge achievement on its own. A great test set for anyone trying to do this: look at the Scott Adams sock puppet controversy on Metafilter [1] and see if you can train something on his public writing to match the "PlannedChaos" commenter's posts and Adams' own tweets. It is probably the closest you could get to a "pure" training set in the sense that presumably Adams didn't think he'd get caught. (And if he did, and therefore did alter his stylometrics, then it's even a better challenge.) [1] http://www.adweek.com/galleycat/scott-adams-caught-defending-himself-anonymously-on-metafilter/28973 http://www.adweek.com/galleycat/scott-adams-caught-defending...