9 ms·
How Not to Sort by Average Rating (2009)
- hood_syntax 9y agoRead this article before and I really liked how to the point it is. More than anything, can I just say how infuriating Amazon's rating system is?
- alexpetralia 9y agoChris Stucchio and Evan Miller have amazing statistics blogs.
- paulgb 9y agoAverages (even with the post's approach) still have the problem of not being "honest" in the game theory sense. For example, if something is rated 4 stars with 100 reviews, a reviewer who believes its true rating should be 3 stars is motivated to give it 1 star because that will move the average rating closer to his desired outcome. A look at rating distributions shows that this is in fact how many people behave. Median ratings are "honest" in this sense, as long as ties are broken arbitrarily rather than by averaging. Math challenge: is there a way of combining the desirable properties mentioned in the post with the property of honesty? I suspect there is but I haven't tried it.
- frgtpsswrdlame 9y ago>A look at rating distributions shows that this is in fact how many people behave. This is really interesting, do you have anything where I could read more about it?
- paulgb 9y agoHere's an analysis from a couple years ago: http://minimaxir.com/2014/06/reviewing-reviews/ http://minimaxir.com/2014/06/reviewing-reviews/ In particular, the conclusion: "The reviews on Amazon’s Electronics products very frequently rate the product 4 or 5 stars, and such reviews are almost always considered helpful. 1-stars are used to signify disapproval, and 2-star and 3-stars reviews have no significant impact at all. If that’s the case, then what’s the point of having a 5 star ranking system at all if the vast majority of reviewers favor the product? Would Amazon benefit if they made review ratings a binary like/dislike?"
- bnegreve 9y agoI remember a talk by Thorsten Joachims at ECML/PKDD 2013 where he was talking about this, you can watch it here. https://www.youtube.com/watch?v=fX9lj0UdB9s https://www.youtube.com/watch?v=fX9lj0UdB9s Around 11:40 he shows evidence of this "dishonest" behavior. As far as I remember, the whole talk was very good. He has some publications on the topic.
- asr 9y agoI am having a hard time digging it up, but I remember reading some reporting on the Netflix Prize that said (before Netflix abandoned the star system) that many users rated things only one-star or five-star. But, counter to the OP's point, I wouldn't assume this is an attempt to move the average; I would guess this is for a number of reasons, including because it's too much mental energy to decide if a product (film) is worth four or five stars, if you rate something you are often just trying to say "liked it" or "didn't like it."
- SamBam 9y agoThat's an interesting hypothesis, but I'd want to see more evidence that "a reviewer who believes its true rating should be 3 stars is motivated to give it 1 star." Surely you can't prove this simply by noticing that there are many 1- and 5-star reviews, as there could be many other reasons for that. One obvious one: people who strongly like or strongly dislike a product are more likely to take the time to review. I personally have never felt the need to review something if I felt "meh" about it. One sample study might be to see how people's ratings change if they have a chance to see the average rating first or not, but that would be a tricky study as you'd need to get people to buy something without seeing the ratings.
- paulgb 9y agoThat's a good alternative hypothesis. It could also be that people's experience with a product really is bimodal: if I order an alarm clock that works as advertised, it is easy to get 5 stars, if it doesn't work at all, it's 1 star. Your explanation works better for why the distribution persists in books and movies though. In any case, I find the mechanism design angle interesting regardless of the behavioral angle :)
- gknoy 9y agoI've often hypothesized that most people are more likely to leave a negative review when they are upset, than a positive review when they like something. It certainly holds in my case: I nearly never review things, because it's a giant hassle, so I have to feel really strongly to be willing to spend the time on a review.
- CM30 9y agoYou're right. It's been proven a bunch of times that people are more likely to leave negative reviews than positive ones, or to remember bad experiences more overall. Zendesk actually did a survey on this back in 2013: http://cdn.zendesk.com/resources/whitepapers/Zendesk_WP_Customer_Service_and_Business_Results.pdf http://cdn.zendesk.com/resources/whitepapers/Zendesk_WP_Cust... And American Express found something similar in their Global Customer Service Barometer survey as well: http://about.americanexpress.com/news/docs/2014x/2014-Global-Customer-Service-Barometer-US.pdf http://about.americanexpress.com/news/docs/2014x/2014-Global...
- nfriedly 9y agoI wonder if a system that assigned weights to each individual user's rating based on that user's rating history could help there - if a user always rates products with 5-stars, then another 5-star rating shouldn't have nearly as much weight as one coming from a user that gives a fairly balanced range of ratings. I'm not sure if that would actually work better in practice, but it's at least an interesting idea.
- fnordian_slip 9y agoAt first your system seemed to me like a solution to bought reviews (mturk and otherwise) and bots. Then I realized that it would just incentivise bots to add 1-star reviews to random products once their creators figure out this mechanism. Sometimes these problems make me sad, it could all be so nice and easy if it weren't for these bad actors.
- nfriedly 9y agoYea, that's the main reason why I'm not sure it could be made to work.
- BearGoesChirp 9y agoIt would be a lot more work, but one could check the validity of ratings based on how other users rate things compared to this user. Say there are 4 games. Most users who rate all 4 rate them similar, except for the 4th game that always gets really low. So 5,5,5,1 is a normal expected rating, but 1,1,1,5 isn't. So 5,5,4,4 from a high rater or 2,2,2,2 from a low rater would be given more weight than a 1,1,1,5. Other things can be added such as weighting a user's ratings a low impact if they have too few scores to determine ratings from. This reminds me of the problem of determining the answer key to a multiple choice test given only the answers of the test takers.
- bllguo 9y agounfortunately sample sizes will probably decrease drastically when you're looking at the subset of people who rated 4 specific games, or other such cases
- hyperpape 9y agoJohn Gruber has been arguing that the only meaningful way to do ratings is a simple thumbs up/thumbs down. I don't necessarily agree, but I see the appeal. I usually don't want ratings, I want the Wirecutter treatment. Sometimes, I know/care enough to really research the topic, in which case star reviews are relatively unhelpful. The rest of the time, I just want someone trustworthy to say "buy this if you want to pay a lot, buy this if you want something cheap, but this third thing is no good at any price".
- jonknee 9y agoEven Netflix finally moved over to up/down and they were famous for squeezing every drop out of their previous star based reviews [1]. In theory stars work better, but the issue seems to be everyone has a different ranking system. For example, Uber seems to think anything but a 5/5 is a failure. I know this so I skew to accommodate, but in my personal ranking system I've only had a couple 5 star rides (someone really going above and beyond). Up/down with an optional qualifier afterwards (e.g. "why were you unhappy?" after a thumbs down) seems to remove a lot of confusion. [1] https://en.wikipedia.org/wiki/Netflix_Prize https://en.wikipedia.org/wiki/Netflix_Prize
- yoz-y 9y agoMaybe the problem is with stars and wording. Currently most of systems are worded (e.g. amazon) in a way that 3 stars is the base and people would add stars if their expectations were exceeded and remove them if they were not met. However at all places that I have seen it is like you say 5 stars is for expectations being met and it only goes downhill from that. I think that a wording and iconography in 4 steps could be useful. -2 = something really bad happened, -1 = below expectations, 0 = happy customer, 1 = exceeded expectations. Forcing people to write a detail on any rating other than 0 would make most ratings 0. Angry people usually like to write comments anyways.
- ghaff 9y agoI'm not sure an expectations-based rating system is the norm though. To use an example I gave elsewhere. I order a cable from Amazon. It works. Therefore it met my expectations of a working cable. Yet, I think most people would interpret a 3-star rating as my being lukewarm on my purchase. I'm not. But what the heck do I expect a cable to do other than being a fair price and to work? With respect to movies. Some movies get really built up and I go in expecting great things (e.g. Fury Road). I come out thinking they were just OK. So maybe 3 stars. But definitely not -1 or 2 stars. My personal expectations aren't necessarily a good baseline.
- thanatropism 9y agoIndividual preferences cannot be aggregated into something that resembles a preference ranking. The most cited formalization of this is Arrow's impossibility theorem, but choice aggregation is this whole theory. _Judgement_ is a slightly different problem. There's an entire issue (#145) of the _Journal of Economic Theory_ on this, but the panorama is still quite bleak, and the reddit approach is far from state-of-the-art. (Personal experience: I've "returned" to reddit (I swore off facebook but I'm still addicted to having something on my phone), and the only way to get people to interact with you is to browse the "new" queue. Once something is "hot" it's basically dead -- new comments are queued to the end even if they're rising fast, and no one replies to you).
- bo1024 9y agoYes, but single-peaked preferences are a special case that apply here, and where the median is truthful. (For those not familiar: single-peaked preferences assumes that the person always prefers the final rating to end up closer to their personal rating. So if I believe the restaurant is 3 stars, I'd most prefer it gets rated actually at 3 stars, and I'd rather see 2 stars than 1 star. If all the raters have single-peaked preferences, then using the median to produce the final rating is truthful: A person can't move the final rating closer to their own belief by lying. The mean is not truthful: If the current average is 4 and my rating is 3, I can pull the average closer to 3 by giving a 1-star review.)
- zolloie 9y agoThe flaw with that line of criticism is that it makes assumptions about the meaning of the ratings. Note, too, that Arrow's impossibility theorem applies to ranking but not ratings. That also applies to a very simplified, idealized case which can be superceded by more sophisticated voting/rating systems.
- svachalek 9y agoThere's also the hostage-taking effect, most notable in the iOS app store. "This 1-star review will be changed when you give me what I want."
- stordoff 9y agoIt also doesn't help when reviews aren't made on the merits of the product. Pretty common on, e.g., Steam: > DOTA 2 users then had the brilliant idea to do the dumbest thing any fanbase can do to a game, flood Metacritic with bad user reviews. The slew of zeros since the forgotten Diretide has dropped the game's user score about two points to a 4.5. https://www.forbes.com/sites/insertcoin/2013/11/02/valve-forgot-diretide-halloween-and-dota-2-fans-are-not-pleased https://www.forbes.com/sites/insertcoin/2013/11/02/valve-for...
- dota 9y agoHalf-Life fans are currently leaving negative reviews on Dota 2 because Valve won't make HL3. Recent reviews went from "overwhelmingly positive" to "mixed". https://arstechnica.com/gaming/2017/08/steam-reviewers-bomb-dota-2-over-lack-of-half-life-3/ https://arstechnica.com/gaming/2017/08/steam-reviewers-bomb-...
- logfromblammo 9y agoValve should probably just farm out HL3 to Obsidian, and continue printing money with Steam.
- pvdebbe 9y agoOpen-world Half-Life does sound intriguing.
- grandalf 9y agoExcellent point. Per your question I'm also curious if there is a way to make aggregate ratings more useful as a quality measurement. For instance, if I see a product on Amazon with a 4.8 average rating but notice a lot of very angry 1 star ratings, I'm likely to infer that there may be quality control problems. Amazon displays a histogram so the shopper can assess the meaning of the distribution heuristically. There's also the issue of whether ratings should be absolute or based on value. If I buy some obviously knockoff ear buds for $6 and they are way better than expected, I'd give them 5 stars, but if they had cost $50 I'd have given a three star review. So for shopping it seems that there are multiple signals being aliased into a single star rating.
- paulgb 9y agoSomewhat related, I think it was Nate Silver who theorized that given enough time all restaurant reviews trend towards four stars. The theory is that if something gets less than four stars it attracts a niche crowd that appreciates it uniquely (and rates highly), while if it gets five stars it attracts a general crowd that doesn't have a particular appreciation (and rates poorly).
- ghaff 9y agoAnd/or restaurants that everyone hates tend not to stay in business. There's also a more general effect that ratings (and reviews) affect the behaviors of people who haven't rated yet. I wouldn't rate many movies that I see below a 3/meh/OK level. That's not so much because I grade inflate but because I actively seek to avoid movies I'd rate 1 or 2 (and, indeed, mostly 3).
- dredmorbius 9y agoSome years back I designed a multi-point rating system for a social media site. I used it precisely as you describe. Ended up in a discussion/argument with one of the users (not aware of my role) over whether or not that constituted "abuse" of the system. It was pointed out (by others) what my relationship to the system design was. I remain amused by the episode.
- bo1024 9y ago> Math challenge: is there a way of combining the desirable properties mentioned in the post with the property of honesty? I suspect there is but I haven't tried it. The suggested method seems to be asking for binary responses (like/dislike), then aggregating them with the confidence-bound formula. This should be truthful in, e.g., a model where users who like it want to maximize the score and users who dislike it want to minimize the score.
- KVFinn 9y ago>Averages (even with the post's approach) still have the problem of not being "honest" in the game theory sense. For example, if something is rated 4 stars with 100 reviews, a reviewer who believes its true rating should be 3 stars is motivated to give it 1 star because that will move the average rating closer to his desired outcome. Also I'd like to know if a 5/10 rating is mostly 5s, or an average of mostly 1s and 10s.
- baddox 9y agoIf it's true that a significant number of people give 1-star reviews to drive down the average (I'm skeptical without seeing evidence), then would people really understand the idea of median ratings and stop doing that?
- panic 9y agoYou can also model each vote as an "agent" that tries its best to move the star rating toward its desired value. If the current rating is a 4, each "agent" with a vote less than 4 will throw a 1 into the average, and each vote greater than 4 will throw a 5. This process converges, though the rating tends strongly toward 3 (or whatever the middle value is).
- paulgb 9y agoGood answer, this has a nice property that it can be applied to any reasonably behaved average-based system to get an honest mechanism. For a plain average it is equivalent to the median.
- ajennings 9y agoNot actually equivalent to the median. If all the scores are 1s and 5s, then the median will be a 1 or a 5, but the average will be somewhere in between. The problem is slightly easier to understand if we consider the grading scale from 0 to 100. Then every agent, trying to manipulate the score as much as they can toward their "ideal grade", will submit a grade of 0 or 100. The average will converge to the unique number, X, where X percent of the graders want the final grade to be above (or equal to) X.
- paulgb 9y agoYes, you're right and I was wrong.
- dragonwriter 9y agoNumeric ratings are GIGO anyway, since cultural differences in how people map satisfaction to star ratings mean that the same number of people with the same degrees of preference for your product can produce a near-infinite array of different sets of preference ranking simply depending on how preferences are distributed among subcultures.
- ajennings 9y agoGood question. Usually we consider "aggregation functions" with a fixed number of graders, N. It has been proven that if you want an aggregation function that is: - anonymous: all graders treated equally - unanimous: if all graders give the same grade, then that must be the output grade - strategy-proof: a grader who submitted a grade higher (lower) than the output grade, if given the chance to change their grade, could do nothing to raise (lower) the output grade - strictly monotone: if all graders raise (lower) their grade, then the output grade must rise (fall) then your aggregation function must be an "order statistic": the median (if N is odd) or some other function which always chooses the Mth highest input grade. If you relax the last criterion to: - weakly monotone: if all graders raise (lower) their grade, then the output grade must rise (fall) or stay the same then your aggregation function must be "the median of the input data and N-1 fixed values". As an example of this last type of function, let's take @panic's idea that each grader has an honest evaluation between 0 and 100 but has an agent that submits a fake grade (0 or 100 usually) to pull the average toward their honest evaluation. As I say in a descendant comment, this system will converge to the unique number, X, such that X percent of the graders want the final grade to be X or above. You noted that this whole system (the average and the agents) is strategy-proof, so each grader should be honest with their agent. We might as well pull the agents into the system and say, "submit your honest evaluation and we'll calculate X, the unique number such that X percent of the graders want the final grade to be X or above." This is an aggregation function. It is anonymous, unanimous, strategy-proof, and weakly monotone. I call it the "linear median" in my PhD thesis. Rob LeGrand called it "AAR DSV" in his thesis. We've been calling it the "chiastic median" more recently. It has some interesting properties. Considered in the context of "the median of the input data and N-1 fixed values", with 100 graders, the 99 fixed values are 1,2,...,99, and this function always returns the median of the input data with these 99 fixed values. (No matter how many graders there are, the fixed values will equally divide the number line between 0 and 100.) You can see chapters 5-8 of my PhD dissertation for more info: http://ajennings.net/dissertation.pdf http://ajennings.net/dissertation.pdf Now, you're thinking about how the grade changes when a new vote is added, so we're really talking about a family of aggregation functions, one for each possible number of graders. We want each one to be strategy-proof in itself, but we also need to consider how they relate to each other. Do you want strict monotonicity or weak? (I find strict monotonicity too restrictive, myself.) If you say "strict", then for each N you need to choose which order statistic you want. If you say "weak", then for each N you need to choose N-1 fixed values and you'll always take the median of the input data and the appropriate array of fixed values. In my thesis (section 7.2) I talk about how you can create a "grading function" to unify a family of aggregation functions, but I don't think that's a perfect fit since we want to somehow "punish" subjects that don't have very many grades (that's what the OP is about). Do we want to pull them towards 0, or pull them towards some global neutral value (like 3 out of 5)?
- pcollins123 9y agoFor every service there are very relevant factors about what makes the product good or bad. If you don't separate these 2 or 3 factors, then over time everything becomes a score of 3.6. In the case of Amazon, the relevant options are: 1. Likert scale of quality: a junk, just don't buy it b cheap and works good enough for occassional use c higher quality: willing to spend more and you'll get a much better outcome. d overpriced 2. bad shipping, bad vendor, poor customer service I hate seeing a bad review for a product based on the last item, they're normally outlier issues or whiners and I normally try to filter them out. In the case of rotten tomatoes it is, again a different set of parameters.
- intenscia 9y agoImplemented this after discovering it via https://www.gamasutra.com/blogs/LarsDoucet/20141006/227162/Fixing_Steams_User_Rating_Charts.php https://www.gamasutra.com/blogs/LarsDoucet/20141006/227162/F... Works amazingly well and so easy to calculate vs say the way IMDb rates things.
- kstenerud 9y agoThat's what's always annoyed me with Amazon's "sort by average rating" setting. I want to see the top 10 or so items by rating to give me a baseline to investigate from, but instead I get page after page of cheap Chinese crap with one 5-star review each from the resident fake reviewer. Worse than useless. Even a simple change like adding a "show only items with a minumum of X reviews" would be a godsend.
- nehushtan 9y agoWhat's crazy is everyone knows Amazon's ranking is crap, except apparently Amazon - and it's been crappy in the same way for 10 years.
- SamBam 9y agoWhile Amazon certainly has a vested interest in getting people to trust the reviews (they care more about people coming back to Amazon again and again than selling any one product), I wonder if they also have a vested interest in keeping a large number of products from multiple vendors available. If Amazon ranked it's items the "proper" way, such that all one-rating products were far from the top, I imagine it would see a lot more clustering of purchases on the most popular version of every product type. All those variations that were not as popular would receive many fewer purchases, and some of those vendors might simply fold. Amazon may have decided that having a larger ecosystem of vendors is worth more than implementing a better rating system. This, presumably, is not in the customer's interest (unless perhaps the "discoverability" of unknown products is on balance worth the risk). Whatever the reason, Amazon certainly knows about other ranking systems, so it has to have made this choice deliberately.
- folli 9y agoFor Amazon (and equally large companies) I usually tempted to put the proverb "don't attribute to malice which is adequately explained by stupidity" on its head. There's definitely a financial reason behind this.
- oconnor663 9y ago
- donatj 9y agoDoes a decent Fortran implementation exist?
- jimktrains2 9y agoDid you read the article?
- jimktrains2 9y agoTo those downvoting me and sibling, parent asked about a SQL implementation originally.
- tom-lord 9y agoIt's written in the article?
- thanatropism 9y agoArguably what Urban Dictionary is doing is to weigh by "net favorability" in some sense and quantity of votes. Quantity of votes correlates to relevance, particularly because UD is meant to represent popular usage.
- onorton 9y agoWould a better idea for UD be like the original Facebook "like" system? So you only vote if you think it's relevant and only popular definitions sit at the top.
- gleenn 9y agoWe actually switched to Wilson score. Doing it later has some weird effects, when you've already have a lot of people typically voting on the first definition, and then suddenly the order gets switched because something has a higher ratio giving it higher confidence. We're honestly not sure it's done anything that great for UD, sometimes simple is just better.
- thanatropism 9y agoYou might want to weigh less controversial (as in abs(upvotes - downvotes) higher. This would be somewhat like Effect Size in science. The Bayesian approach would be to assume the true vote distribution is binomial and use a beta prior (possibly with Jeffrey's degenerate bimodal prior). Then as the total number of votes increases the posterior distribution tightens. Ranking score is prob(score>0).
- ignawin 9y agoAny blog posts/papers on what the best general approach to onliene reviews is?
- dredmorbius 9y agoGood, qualified, honest reviewers. Hal Varian (UC Berkeley) has some 1990s refs, which remain good. "Grouplens" is the project/product. Randy Farmer literally wrote the book on the topic. There's a book, blog, and wiki. Frankly, Farmer's work, good as it is, largely reinforces my view that Varian captured the essence of the problem, which I've summarised in my opening 'graph. You cannot algorithmically correct for crap quality assessment. If you're interested in the long-form answer, the fields are epistemology (philosophy) and epistemics (science). Enjoy! http://people.ischool.berkeley.edu/~hal/Papers/publish.html http://people.ischool.berkeley.edu/~hal/Papers/publish.html http://people.ischool.berkeley.edu/~ngood/ http://people.ischool.berkeley.edu/~ngood/ http://people.ischool.berkeley.edu/~hal/Papers/japan/ http://people.ischool.berkeley.edu/~hal/Papers/japan/ http://buildingreputation.com http://buildingreputation.com
- jules 9y agoI wrote a blog post about a simpler and statistically grounded method: http://julesjacobs.github.io/2015/08/17/bayesian-scoring-of-ratings.html http://julesjacobs.github.io/2015/08/17/bayesian-scoring-of-...
- eeZah7Ux 9y agoThis is computationally very heavy, but, more importantly, for practical purposes you want to have a tunable parameter to balance between sorting by pure rating average and sorting by pure popularity. Often you also want to give a configurable advantage or handicap to new entries.
- quantdev 9y agoFor a fixed confidence level, it looks computationally light weight: a dozen or so multiplications and divisions plus one square root, which could be approximated if needed. There is no inverse normal needed at run time.
- dperfect 9y agoWhat's the best way to apply the suggested solution to a numeric 5-star rating system (the author mentions Amazon's 5-star system using the wrong approach, yet the solution is specific to a rating system of binary positive/negative ratings)? I suppose one could arbitrarily assign ratings above a certain threshold to "positive" and those below to "negative", and use the same algorithm, but I imagine there's probably a similar algorithm that works directly on numeric ratings. Anyone know? Or if you must convert the numeric ratings to positive/negative, how does one find the best cutoff value?
- amrrs 9y agoWhat we do with 5-star rating system is completely ignore 2,3,4 stars which in a lot of ways just skew our analysis, hence ending up with a new score similar to Nps (5-star minus 1-star) / total stars
- overcast 9y agoWhy bother have a 5-star rating system then? Sounds like Netflix went down the right path, with thumbs up or thumbs down. People either zero it out, or give it five stars.
- dbaupp 9y agoThe author has also written http://www.evanmiller.org/ranking-items-with-star-ratings.html http://www.evanmiller.org/ranking-items-with-star-ratings.ht... .
- poorman 9y agoI reference this article constantly at Untappd. When we were building the NextGlass app, I took much of this into consideration for giving wine and beer recommendations. We recently ran the query on the Untappd database of 500 million checkins and it yielded some interesting results. The "whales" (rare beers) bubbled to the top. I assume this is because users who have to trade and hunt down rare beers are less likely to rate them lower. The movie industry doesn't have to worry about users rating "rare movies", but I would think Amazon might have the same issue with rare products.
- chris_va 9y agoThere is an interesting phenomenon of exclusivity/sunk-cost boosting ratings for rarer or harder to acquire items. That is also a problem with movie ratings (I just noticed that you mentioned movies). Critics (and audiences) at pre-screenings are generally significantly more favorable to a movie than an equivalent group in a normal theater. I would not be surprised if the same thing applied to foreign movies, and other types of "whales".
- infomofo 9y agoUntappd is also weird because you know that the producers of some of these small brewery beers actually look at these checkins. A lot of the beer drinkers I know will prefer to not rate a beer instead of giving it a sub-3 rating.
- amelius 9y ago> What we want to ask is: Given the ratings I have, there is a 95% chance that the “real” fraction of positive ratings is at least what? Wilson gives the answer. Well, you can't answer that question without making assumptions. And these seem to be missing in the article.
- jbochi 9y agoIt's very common to see a "Most Popular" section in a website, but the way it's usually done is not optimized for clicks. Inspired by Evan's post, I wrote "How Not to Sort by Popularity" a few weeks ago: https://medium.com/@jbochi/how-not-to-sort-by-popularity-92745397a7ae https://medium.com/@jbochi/how-not-to-sort-by-popularity-927...
- loisaidasam 9y agoHere's a SO post w/ a python implementation: https://stackoverflow.com/questions/10029588/python-implementation-of-the-wilson-score-interval/45965534 https://stackoverflow.com/questions/10029588/python-implemen... The accepted answer uses a hard-coded z-value. In the event that you want a dynamic z-value like the ruby solution offers, I just submitted the following solution: https://stackoverflow.com/questions/10029588/python-implementation-of-the-wilson-score-interval/45965534#45965534 https://stackoverflow.com/questions/10029588/python-implemen...
- larkeith 9y agoThis article is useful, but the author's tone really rubs me the wrong way - to the point I'm dubious about trusting the information without further sources. Cutting the entire first part ("not calculating the average is not how to calculate the average") would help, as would more accurately titling the piece - no matter how effective this method is, it is NOT sorting by average, strictly speaking.
- toniprada 9y agoOther approach for non binary ratings is to use the true Bayesian estimate, which uses all the platform ratings as the prior probability. This is what IMBD uses in its Top 250: "The following formula is used to calculate the Top Rated 250 titles. This formula provides a true 'Bayesian estimate', which takes into account the number of votes each title has received, minimum votes required to be on the list, and the mean vote for all titles: weighted rating (WR) = (v ÷ (v+m)) × R + (m ÷ (v+m)) × C Where: R = average for the movie (mean) = (Rating) v = number of votes for the movie = (votes) m = minimum votes required to be listed in the Top 250 C = the mean vote across the whole report" http://www.imdb.com/help/show_leaf?votestopfaq&pf_rd_m=A2FGELUUNOQJNL&pf_rd_p=2398042102&pf_rd_r=0XCZT7P6XFKP334FP3GT&pf_rd_s=center-1&pf_rd_t=15506&pf_rd_i=top&ref_=chttp_faq http://www.imdb.com/help/show_leaf?votestopfaq&pf_rd_m=A2FGE...
- kuharich 9y agoPrevious discussions: http://news.ycombinator.com/item?id=1218951 http://news.ycombinator.com/item?id=1218951, http://news.ycombinator.com/item?id=3792627 http://news.ycombinator.com/item?id=3792627, https://news.ycombinator.com/item?id=9855784 https://news.ycombinator.com/item?id=9855784
- gesman 9y agoI think ratings need to be normalized to personal beliefs and preferences of the viewer. In other words - I can care less how Joe Blow rated the product - but it's important to me how likeminded people like me rated the product. Also - Amazon is not making mistake in ratings. Amazon is less interested in selling you relevant product for you. Amazon is more interested to boost it's bottom line, move stalled inventory or move higher margin inventory.
- Animats 9y agoMandatory XKCD: https://xkcd.com/937/ https://xkcd.com/937/
- phunge 9y agoClassic post! This post is like a gentle gateway to the world of Bayesian statistics -- check out Cameron Davidson Pilon's free book if you want to go deeper.
- tabtab 9y agoWhat about having a scaling factor to adjust the impact of quantity (total) of individual ratings as needed? Rough draft: sort_score = (pos / total) + (W * log(total)) Here, W is the weighting (scaling) factor. Total = positive + negative
- autarch 9y agoIMDB uses something like this. It's called a "weighted" rating system. In the IMDB version what happens is that you calculate the average of all ratings of all items, and then push an item's rating towards the average. The fewer ratings it has the more it's pushed. See http://www.imdb.com/help/show_leaf?votes http://www.imdb.com/help/show_leaf?votes for details.
- jules 9y agoThat formula will give an item a higher score the more down votes it gets. A better approach is score = (pos + a) / (tot + b). Where a<b, e.g. a=1, b=2. See this post why that formula follows from Bayesian reasoning: http://julesjacobs.github.io/2015/08/17/bayesian-scoring-of-ratings.html http://julesjacobs.github.io/2015/08/17/bayesian-scoring-of-...
- tabtab 9y agoI'm not sure what you mean in the 1st sentence. Example? The problem with the 2 weights is that it's 2 values that have to be given, and for large quantities neither makes much difference. It's why I used log().
- jules 9y agoThe log(total) term increases without bound whereas the pos/tot term is at most 1, so in the limit of a lot of votes you will beat an item with fewer votes even if all your votes are downvotes. That there are two configurable parameters is a good thing. One parameter controls how much of a penalty you get for having few votes, the other controls how many votes count as "few".
- agentgt 9y agoThis sort of reminds of "voting theory" and if I recall it was proven by I think a nobel prize winner that there cannot be a fair winner. Obviously it's not entirely analogous but I would not be surprised if it mapped over to this domain. Edit: on mobile so late on the link to Kenneth Arrow https://en.m.wikipedia.org/wiki/Arrow%27s_impossibility_theorem https://en.m.wikipedia.org/wiki/Arrow%27s_impossibility_theo...
- petters 9y agoThat theorem, being about when every user provides a complete ranking, does not apply in this case.
- bradbeattie 9y agoI think this article is missing the next step: collaborative filtering. I only care about the ratings it received from people that rate thing like I do.
- alexvay 9y agoI think the article is missing something visual to demonstrate the actual scoring at work. I've made a simple plot in Excel here: http://i.imgur.com/adjaLQ9.png http://i.imgur.com/adjaLQ9.png The number of up-votes remains the same, while down-votes increases linearly. The scoring declining line in grey is the score.
- bwaxxlo 9y agoLabel the axises please.
- shaftway 9y agoHere's a 3d graph showing this as a function of upvotes and downvotes. I think it's clearest with x: [0, 100] y: [0, 100] z: [0, 1] https://www.google.com/search?q=graph+((x+%2B+1.9208)+%2F+(x+%2B+y)+-+1.96+*+(((x+*+y)+%2F+(x+%2B+y)+%2B+0.9604)+%5E+0.5)+%2F+(x+%2B+y))+%2F+(1+%2B+3.8416+%2F+(x+%2B+y))&oq=graph+((x+%2B+1.9208)+%2F+(x+%2B+y)+-+1.96+*+(((x+*+y)+%2F+(x+%2B+y)+%2B+0.9604)+%5E+0.5)+%2F+(x+%2B+y))+%2F+(1+%2B+3.8416+%2F+(x+%2B+y) https://www.google.com/search?q=graph+((x+%2B+1.9208)+%2F+(x...)
- autokad 9y agoa gamma poison might more accurately calculate the rating based off uncertainty of the data