3 ms·
Statistics.... No, absolute ground truth is impossible for this problem. There has been some research that estimated the percentage of fake / spam reviews at 2
by dfraser992 9y ago
Statistics.... No, absolute ground truth is impossible for this problem. There has been some research that estimated the percentage of fake / spam reviews at 2 - 6% but that was from a few years ago. I'm sure the percentage has increased greatly since then [great, another dissertation... My MSc dissertation was on this topic]
Text-only based features are not fabulous, just useful (like 75% correct, using a dataset where ground truth is known). So there is / can be linguistic differences in the writing that indicate the writer is not 'truthful'. Sentiment analysis however, by itself, is abysmal - and the behavior of spammers has changed over the years so they are more knowledgeable how to craft reviews.
But other signals, like relationships to other reviewers, IP addresses, submission time of review... they have been shown to be more accurate - but not in the 90+%. Establishing who is in a spamming group seems to be reliable though, so once that is known, you can more confidently label their reviews as spam [but not 100%, of course, to be fair and objective]
I guess Fakespot takes a stab at estimating correctness and hopes the false negative rate is acceptably low. Yelp OTOH cranks things up so the false positive rate is high....
- forapurpose 9y agoThanks; that kind of contribution is why I read HN. > using a dataset where ground truth is known How is this dataset created?