4 ms·
Agree. Rather than release the code, release the testing data set and/or test methodology used to make a claim of 95% accuracy; no one is harmed and it should g
by firebones 10y ago
Agree. Rather than release the code, release the testing data set and/or test methodology used to make a claim of 95% accuracy; no one is harmed and it should give a good sense of the claim. I wouldn't be shocked to find that the score was achieved by overtraining on a very small and narrow data set (small number of identities), and that the model isn't generalizable.
That said, identification may be a more tractable problem if you have a limited population, additional metadata for features, and normalized writing samples (comparing anonymous reviews to identified reviews, or within community posts, as opposed to trying to compare a set of anonymous tweets to an identifiable dissertation).