10 ms·
Brewing a Better Rating System
- stuartjmoore 17y agoIn this context, it looks like the user has incentive to rate (to get better suggestions), but on rating in general: Why even ask people how they feel? Depending on the content, you can analysis how they use it to get a much more accurate rating. For video: Did they watch the entire thing? Did they leave after a few seconds? Did they share it somehow? That (slightly off-topic) being said, this looks great.
- TrevorJ 17y agoI feel like this really combines the best of the granular 4 star systems with the specificity of a percentage rating. Really good stuff, I'd be interested to hear a follow up with user feedback on this approach and how it holds up long term.
- alabut 17y agoI love that the sliding meter shows tick marks for your previous rankings of other teas - the UI reflects that your judgement of a particular tea is relative to your other experiences.
- dylanz 17y agoExactly. I think the tick marks are the killer feature here. Great idea!
- marcus 17y agoAn interesting idea might be to let the user modify a tick mark in retrospect when scoring a new tea.
- pbhjpbhj 17y agoI suspect under closer analysis ones scoring breaks down to be inconsistent - "in retrospect I like tea X better than Y but not as much as Z, but I rated Z lower than Y because I didn't like it as much as P which had a higher rating", if you follow.
- snprbob86 17y agoIt seems that, pairwise, it is pretty easy to decide. Maybe one could go further than this and eliminate the absolute scale all together (at least at rating time). I'm imaging a UI which asks you to pick a favorite among the item you are viewing and one similar item. You could stop there, or ask repeatedly with new comparison items until the viewed item's position on the absolute scale is unambiguous. The user could provide some rating data with just a single binary decision, but some ajax-y fade out/in of another pair could enable further ratings if they desired.
- m_eiman 17y agoIf you do that, the Elo rating system is a good place to start algorithm-wise. http://en.wikipedia.org/wiki/Elo_rating_system http://en.wikipedia.org/wiki/Elo_rating_system
- jurjenh 17y agoTaking this idea further, one could add multiple orthogonal axes (eg sweetness, bitterness, after-taste etc). Then you could rate each tea against others on each axis - either in a star-slider, scatter plot or on several individual sliders. This would allow you to rate teas against each other based on several aspects, and possibly allow recommendations based on how other people have rated teas - eg 'I want a tea that is not-too-sweet, a little bitter with a lingering aftertaste' Then again, it does add more features / visual clutter and possibly complicates things for people...
- mpotter 17y agoHi, I'm Mike from Steepster. We thought we'd share our new ratings system we just deployed with HN as we think it's relevant for products with customer reviews, ratings, etc. It's our attempt to combat the 4.3 dilemma (discussed here recently: http://news.ycombinator.com/item?id=883890 http://news.ycombinator.com/item?id=883890). Background: Steepster is a community site for tea drinkers to share their tasting notes, get recommendations, and discover new teas. Feedback appreciated!
- cninja 17y agoVery clever. Have you considered making the slider non-linear (the distance on the slider between Yuck and Meh is smaller than between Good and Awesome)? If most people are going to rate their tea somewhere between Good and Awesome, it allows more of the slider to be used. Have you received enough ratings with this new interface to know if my assumption of ratings being clustered is accurate?
- mpotter 17y agoWe haven't considered making it non-linear. It's an interesting suggestion, though I'd hesitate to go that way only because the user then lacks a clear 1:1 model of how the slider directly affects their rating (without explanation). We haven't received enough ratings yet to prove your assumption. When we do and if it does hold true, our assumption now is that we have _enough_ of a scale to expose meaningful differences. You've given us good food for thought, thanks. My general feeling now is that I think it's important to leave the negative portion of the slider intact (however less it's used) to maintain a solid mental model. Might be something to test down the road though.
- callmeed 17y agoMike, great job on this. Very informative. I have 2 questions for you: 1. is your slider from the jQuery UI or other js framework? 2. in regards to combating the 4.3 dilemma, have you found the average ratings on steepster to be lower? maybe its too early to tell, but I'd love to see some sort of curve on your ratings distribution in a future post ... thanks
- mkinsella 17y agoThis is THE best implementation of a ratings system I've seen. Very good job.
- deleted 17y ago[deleted]
- ErrantX 17y agoThe genius is adding some previous scores. I always struggle to rate stuff fairly without anything obvious to compare it with.
- fuzzythinker 17y agoI think main reason sliders aren't used is that users find it too troublesome, hence up/down and 5 stars are mainly used. I remember from my pys class that a 7 point rating system is best. But the 5 stars' simplicity and ubiquity probably trumps the benefits gained by a 7 point system. I think the best compromised is a 5 star UI implemented as 6 points by allowing 0 point assignments.
- fuzzythinker 17y agoFor those who marked me down, would you please comment on reason? I'm getting tired of spending my time commenting and getting disapproval without reason. I don't think down votes should be on disagreements; it should be on spamish, childish, or comments that does not add anything to the topic. My main point is that sliders aren't used much because they are too troublesome for a typical user. If you disagree with that, please add your opinion. I'm not trying to take anything away from the author. In fact, I think it's an ingenious idea. But I usually dislike repeating "wow, cool" comments since so many others have done so already. It's part of my DRYness kicking in.
- nkurz 17y agoI voted you down because you asserted that a system was 'best' because you remember from a 'pys' class. This has to be one of the weakest 'arguments from authority' I have seen. You then asserted that a 5-star allowing zero is even better. Then why didn't your (psychology?) professor say so? I didn't vote down because I disagree, but because you haven't made much of a case. I also downvote the 'wow, cool' comments as unhelpful, and upvote the comments that seem like they will lead to useful discussion. Without intending offense, I didn't think your comment was pitched at the right level for this audience. Personally, I think you are on the right track, although I think 5 stars allowing halves is even better. Interestingly, Netflix (experts in this field) started out with allowing half-stars and then got rid of them, making me worry that they know better than I.
- fuzzythinker 17y agoYou are taking every single word of my comment too seriously. If every assertion needs to have strong backing in order to be commented, the hn comments will probably be only < 10% of what it is now (again, just a guesstimate, don't take this one too seriously too). I forgot if my professor has research backing for a 7 point sys being "best", maybe he did, maybe he didn't. But I don't think I need to remember if there was indeed research backing for it to add to the discussion. Again, I don't think you should down vote every discussion just because they didn't state the research backing, but I'm not the one to tell you that, maybe others can comment on this. As for the 7 point system being "best" (for general purpose rating), I remember it's because 5 star does not give enough granularity, while 10 points is too much. Maybe that's why Netflix took that out. Now why not 6, 8, or 9? I forgot, again, maybe there was research being done. As for my "idea" of allow a 0 on a 5 point system; it makes it a 6 point system while retaining a 5 point UI that everyone is accustom to. What is wrong with that? Again, just asking for discussion, not trying to say it IS the best. Now back to the topic of down vote because I don't have enough backing. If I need backing in order to comment, I wouldn't even be able to comment any of this. Is this what you think is the way hn should work? Also, in order to not make you think I have the research to back up my thoughts, I need to say that in almost every sentence. I also don't think that should be the way hn works.
- lonestar 17y agoThe problem with this system is in the sorting. The list of "Highest Rated" teas is dominated by results where 1 person rated the tea 100. Steepster should use a Bayesian average (http://en.wikipedia.org/wiki/Bayesian_average http://en.wikipedia.org/wiki/Bayesian_average) so that the uncertainty of a small number of ratings is reflected in the sorting.
- mpotter 17y agoYeah, sorting is an issue we're still looking at (and is still very much in transition considering the new rating system). Appreciate the suggestion! We'll add it to our list of potential solutions.
- selven 17y agoStart everything off with a single 50-point score. That way one person ranking it 100 will bump it up to 75, the next to 83, and so on. Such a system would cause teas that have more people upvoting them to rank higher than those that just happen to have one or two good opinions.
- Eliezer 17y agoThat's a special case of a particular sort of Bayesian average.
- thinksketch 17y agoThis is very cool thank you. I posted earlier today about the need for a better rating system than the five star system. I'm really glad to see you working on a great solution. Thanks!
- bhellman1 17y agoAbout time someone created a better rating systems. (Stars RIP). It will be interesting to see if users like and use the slider.
- zeeone 17y agoMeh...
- robryan 17y agoSomething else you could think about in a rating system like this would to instead of using generic faces, you could associate each with a common tea that most tea lovers have tried. The notches kind of do this but theres always the risk of someone rating there first tea 80, then deciding subsequent teas after are better so they need to be rated higher, when the first one should have been more around 60.
- mhartl 17y agoThis is cool, but I think virtually all rating systems suffer from the same basic problem: there's no way to turn it up to 11. Take movies, for example. They are usually rated on a four-star scale. And yet, a three-star movie is a clear success. Few movies can realistically aspire to more than three stars. Even many four-star movies are really just trying desperately to avoid two-star land. Francis Ford Coppola was sure he was going to be fired any day from The Godfather. The production crew and actors on Star Wars thought it was practically a joke. Please, God, let Star Wars not be a B movie, they must have been thinking. When you say ★★★ out of ★★★★, you make it look like it wasn't good enough: 75%. Movies really should be rated on a three-star scale: ★★★ out of ★★★; ★★★ = A = 100%. Anything else is gravy. So, rate tea on a three-star scale. Three stars means "excellent tea, no clear way to make it better". ★★★½ means "Whoa, there is something better than ★★★!" ★★★★ means "This is The Godfather of tea! This tea makes me an offer I can't refuse."