3 ms·
Wait, is that true? From the article, it does seem that the training and test sets were both generating from a dubious data collection process: scraping URLs t
by throwawayjava 9y ago
Wait, is that true?
From the article, it does seem that the training and test sets were both generating from a dubious data collection process: scraping URLs they knew were fake/satire/real and then applying a label based upon the domain.
But, they had more features than just "source" in their model. so while it's possible their data collection method means that their model is way over-fit and basically just tests of proxies of "source", it's not prime facie obvious that this is the case...? Or am I missing something?
- minimaxir 9y agoThe website lists domain name as a model feature: https://machinebox.io/docs/fakebox https://machinebox.io/docs/fakebox > Domain name - Some domains are known for hosting certain types of content, Fakebox knows about the most popular sites Fake news sites typically register bespoke domains, though.
- throwawayjava 9y agoOh, I misinterpreted the article. If the test data was collected in the same way that the training data was collected, this was a rather ridiculously round-about way of writing what should've been a 5 line perl script, and I'm really curious where the 5% error could've possibly even come from :-\