5 ms·
The problem with that approach is that it is training the classifier to solve a different problem than what is claimed. It assumes that there is no difference b
by firebones 10y ago
The problem with that approach is that it is training the classifier to solve a different problem than what is claimed. It assumes that there is no difference between how people write when they are identifiable and when they are not (since it would only train on identifiable samples that are anonymized after the fact). Further, it would require solving the problem of identifying authors between different media--which would be a huge achievement on its own.
A great test set for anyone trying to do this: look at the Scott Adams sock puppet controversy on Metafilter [1] and see if you can train something on his public writing to match the "PlannedChaos" commenter's posts and Adams' own tweets. It is probably the closest you could get to a "pure" training set in the sense that presumably Adams didn't think he'd get caught. (And if he did, and therefore did alter his stylometrics, then
it's even a better challenge.)
[1] http://www.adweek.com/galleycat/scott-adams-caught-defending-himself-anonymously-on-metafilter/28973 http://www.adweek.com/galleycat/scott-adams-caught-defending...