3 ms·
I did my dissertation on this very topic (coming soon to my website... at some point). I do not know what tactics or systems Amazon has that deal with fake revi
by dfraser992 9y ago
I did my dissertation on this very topic (coming soon to my website... at some point). I do not know what tactics or systems Amazon has that deal with fake reviews, or if they even really bother beyond the easy wins like tracking IPs. My dissertation was on textual analysis, which is the hard part - the other signals like IP and behavioral related ones (like # of reviews posted in one day) are more fruitful.
The heuristic you seem to be using is a logical one, but requires more data analysis than Amazon might be willing to put effort into. Looking at these reviews, they are... well, so many on the same day is suspicious. My first theory is that some fake review writing company got tasked with flooding Amazon, and so reviews got farmed out to writers. Or someone has invented a GAN that writes good reviews, or at least a good first draft. I'd have to analyze the data to have more of an opinion. But yeah, verified purchaser means little.
The real question is, how economical it is for Amazon to really care about fake reviews? Buyer beware etc and there are more than enough scams running on Amazon / EBay / etc that I'm sure they're just treading water all the time. They have sued review writing outfits, so they care somewhat, but only after the problem got written up in enough newspapers... It is a hard job, trying to analyze all the data coming into their systems each day. I'm not sure any company has really implemented a lot of the research I read about in my literature survey.
ReviewMeta is another site I'd trust:
https://reviewmeta.com/blog/faq/ https://reviewmeta.com/blog/faq/
- bespoke_engnr 9y agoThis is really interesting, thanks. What's your website? I'd love to see that dissertation when it's up. Fascinating stuff.
- dfraser992 9y agohttps://douglas-fraser.com/datadata/ https://douglas-fraser.com/datadata/ this will be the blog (at some point). The overall idea of the dissertation was to see if combining different ways of processing the text of the reviews (classifiers using features from analysis of the grammar, vocabulary, etc) into a custom heterogeneous ensemble was better than using one classifier and the traditional ensemble creation methods (AdaBoost, bagging, etc). I figured creating a more holistic view of the text would be better; other studies have done this, but not to the extent I did. And I analyzed exactly why things did or did not work. So it was just fundamentally a exercise in NLP; I did not use other signals like the # of reviews submitted in one day or other things like that. My gut says this general idea (a more holistic view) would apply to classifying other text, like fake news. But proving that is yet another project. I still have a couple more angles (dependency and constituency parsing, framing) to add to the mix, so I'm not totally done. It will be a long series of blog articles. And I ended up having to deal with the problem of diversity vs. accuracy, so the dissertation went down a side road. My supervisor said it could be two potential papers for publication instead of one... At least I won't be bored for the next year. Thanks for your interest! If you send me your email (dfraser@... is mine), I can send you the PDF, or pointers to other info about the research into fake reviews in general (e.g. using other signals like # of reviews/day); I'm not going to get the blog up soon - already dealing with a ML project for Network Rail here in the UK.