3 ms·
Original author here. For the academically inclined, there is a critique of this approach in this paper: http://www.dcs.bbk.ac.uk/~dell/publications/dellzhang_
by EvanMiller 15y ago
Original author here. For the academically inclined, there is a critique of this approach in this paper:
http://www.dcs.bbk.ac.uk/~dell/publications/dellzhang_ictir2011.pdf http://www.dcs.bbk.ac.uk/~dell/publications/dellzhang_ictir2...
Of course, I think the authors miss the point of the algorithm, since I basically wanted a system that is one-sided (i.e. false negatives are OK but false positives are bad).
Also, if you deal with more than two outcomes you might be interested in multinomial confidence intervals, described here:
http://www.math.wsu.edu/faculty/genz/papers/mvnsing/node8.html http://www.math.wsu.edu/faculty/genz/papers/mvnsing/node8.ht...
The application to 5-star systems is not straightforward, since it's not clear to me how stars relate to each other. Is it a linear scale? Are they discrete buckets? Or maybe we want to use Tukey's froots and flogs? I'm not sure.
By the way, I'm coming out with a stats app for Mac soon that implements this algorithm and much more. Drop me your email address if interested:
http://wizard.evanmiller.org/ http://wizard.evanmiller.org/
- NathanRice 15y agoI appreciate people who take the time to apply math to things in the real world, and share it with non academic crowds. Thanks for that. 5 star rating systems are obnoxious. From a mathematical perspective, if you treat them in an ordinal fashion they are poorly behaved, and if you treat them categorically, you lose the relationship between stars. There seems to be some popular movement towards binary rating systems, and I think that is great. Not only do people tend towards binary rating behavior in the real world (only rating a movie they thought was very good or very bad) but they admit a much cleaner mathematical treatment.
- tripzilch 15y ago> 5 star rating systems are obnoxious. From a mathematical perspective, if you treat them in an ordinal fashion they are poorly behaved, and if you treat them categorically, you lose the relationship between stars. Helping out a friend with a statistics test, I recently read up about the Wilcoxon Signed Rank Test[1]. Now this one is intended to get a p-value for experiments with "before" and "after" measurements, but what I got the idea it's trying to do, is to use the ranks of a not-very-normal behaving random variable, turn that into a summation of lots of things, so due to the central limit theorem you can treat it as a normal distribution again. Though thinking about it, in this case it's the rank we're after, so maybe it's not useful at all. But it gives an interesting idea about the tricks you can pull if your input data isn't quite the sort of type that you can analyse very well. [1] http://en.wikipedia.org/wiki/Wilcoxon_signed-rank_test http://en.wikipedia.org/wiki/Wilcoxon_signed-rank_test
- deleted 15y ago[deleted]
- plainOldText 15y agoI've just seen a screenshot on youtube. It looks interesting.
- YokoZar 15y agoI really appreciated the post. Unfortunately we're using 5 stars, and need to do it a bit less wrong. The main thorniness of 5 stars is that you have to answer the question of what the difference in star ratings actually mean. Is going from one star to two stars the same as four to five? Probably not, based on the way users rate, which means an algorithm like the arithmetic mean that treats them the same is wrong. Personally, I think it's very reasonable to treat rating stars the way we should treat grades: as ordinal data, where we know that a higher rating is better but assume nothing beyond that. The difference between an F and a D is not the same as an A and a B, and the same is likely true of 1-2 stars and 4-5 stars. I have made an attempt at applying this idea to the Ubuntu App Store's rating algorithm. I'm very much interested in comments. https://bugs.launchpad.net/ubuntu/+source/software-center/+bug/894468 https://bugs.launchpad.net/ubuntu/+source/software-center/+b...
- Someone 15y ago"we know that a higher rating is better." Typically, you do not know that, especially if you are comparing across reviewers. Some will only score zero and 5 stars, other will have 10% two stars, 80% three stars, and 10% four stars, yet others will have 10% three stars, 80% four stars, and 10% five stars. If you have sufficient data (rare), it may be possible to (somewhat) correct for that. IIRC, this was something that helped in winning the Netflix challenge. For example, http://www.netflixprize.com/assets/GrandPrize2009_BPC_PragmaticTheory.pdf http://www.netflixprize.com/assets/GrandPrize2009_BPC_Pragma... models the assignment of stars as two parts: - modeling the user's appreciation of the movie - modeling how the user translates his appreciation to a star rating "where we know that a higher rating is better but assume nothing beyond that" Problem with that is that you throw away information with that assumption. You do know that 2 star scores are very unlikely to be about very good items.
- alexchamberlain 15y agoMathematician not a statistician... Would it be reasonable for 5 stars to normalise the data? Should star ratings be on some distribution, for instance?
- NathanRice 15y agoIn the binary case the usual treatment in statistics is to use the logistic function (and logit) to work with real numbers, then transform back into probability space as the last step. This is a little flakey for ordinal numbers, and the usual treatment is to use a learning algorithm to find a mapping from real numbers to ordinal values, either explicitly (if you need a "score") or implicitly. Support vector machines, radial basis functions and neural networks are typically used.
- alexchamberlain 15y agoAre we going to see an Nginx module?