5 ms·
The author didn't benchmark to see if this system is actually any better at predicting outcomes than vanilla Elo. That's how you determine if your implied win p
by defertoreptar 6y ago
The author didn't benchmark to see if this system is actually any better at predicting outcomes than vanilla Elo. That's how you determine if your implied win probabilities are accurately being derived from rating differences. The author seems to be under the impression that there's something fixed and concrete about an 1800 rating, but when you change the system, you also change what an 1800 rating means in the first place.
Some of these complaints are solved by existing systems, namely Glicko. For example, rating deviation helps with experienced players (low RD) losing points to newer players (high RD). It also has a built-in way to discourage inactivity. Players' RD increase over periods of inactivity, so they can be excluded from the leaderboard after reaching a certain point. That allows us to maintain their rating without decreasing it. After all, that's our best guess of the player's skill. It's just a less reliable guess over time.
- BSTRhino 6y agoThere have been four rating systems, including Glicko and TruSkill, received lots and lots of complaints for both those systems. This new system receives few complaints. Tested across 135000 players. If the players had not complained so much, we would still be on Glicko. Those are the facts. The theories as to why that is are up to you.
- jsnell 6y agoOptimizing a rating system for minimal complains, maximum player engagement, or some similar metric is of course totally valid. It reminds me of Sirlin's story of being hired to design a rating system for Starcraft 2, and optimizing for totally different things that Blizzard wanted [0]. If the author (you?) had just described it in those terms, it'd be hard to object. But the article goes further, and makes claims about the system being more accurate due to a different rating curve. That's the claim that would need to be justified by actually comparing whether the predictions the new system makes really are better. [0] http://sirlingames.squarespace.com/blog/2010/7/24/analyzing-starcraft-2s-ranking-system.html http://sirlingames.squarespace.com/blog/2010/7/24/analyzing-...
- an_opabinia 6y agoPart of a matchmaking algorithm that increases user engagement is telling a story about how it's more fair though. Like our tolerance for losing is acquired. Most normal people losing in League for the first time stop playing, usually forever. Just randomly visit your friend's match histories in League, frequent players have many days of long losing streaks. If you're just conditioned to play despite losing, great, in a Darwinian way (surviving, being around to be measured) you will be representative of the average player in League. And there are so many League players with such long retention you cannot possibly argue that skill-based matchmaking is the core component of user engagement. His dataset is interesting because it will necessarily overrepresent people who kept playing despite the old system. That sort of refutes its importance - I mean sure people complain but they keep playing, so was it really that important? So what if complaints go down? Those are important goals, and also, it's still an interesting twist in multiplayer game design. You just gotta interpret it as a commentary on a whole system even if it doesn't narrowly talk about a scientific objective like performance prediction.
- BSTRhino 6y agoPlayer retention was significantly lower before the new rating system was introduced. It wasn't just complaints, it's just that is a more direct metric because multiple changes may have affected the retention.
- Natsu 6y agoWhat you said makes me wonder about a totally different way of using a metric of how likely the person is to win a given match. That is, perhaps the system could be engineered to maintain a more even win/loss ratio so that people don't go on super-long win (or loss) streaks in general by adjusting who they get matched with. It probably wouldn't work that well towards the edges, but around the middle it might work well enough.
- teen 6y agoThis has what dota 2 does. Every win / loss is worth the same points, the games are extremely close in terms of player rank, so and it's a team game so you get more variance than just your individual skill. Players eventually settle close to 50 percent win rate
- cortesoft 6y agoYou have to decide what the purpose of the rating system is. Using it as a reward system for players to feel accomplishment is a different use case then trying to correctly predict the likely outcome of a game. Personally, if I was designing a rating system, I would use two separate systems. One would be like the one in this game; publicly viewable, pleases the players, and gives a sense of accomplishment. Then, I would have a second, internal only rating system that players can't see but is used for matchmaking to make sure people are matched up to players with as close to equivalent skill as possible.
- nobodyandproud 6y agoSome games do this already, and players are very unpleased by the results.
- afterburner 6y agoProbably because they get matched to tougher opponents if they're better?
- wanderer2323 6y agoIt's kinda disheartening being matched to plat players in your silver promos.
- dmos62 6y agoI find similar solutions experientially two-faced and frustrating. Imagine the matchmaking engine was a person behind a desk that you interact with. If he consistently told you one thing and then did something else you'd be displeased.
- ajuc 6y agoThat's what xp is in sc2. For several years xp was the only number that was visible - your true mmr was hidden. Peple knew this and ignored xp altogether.
- mcnamaratw 6y agoMy understanding was that the system consists of using the historical odds of winning (given the rating difference). If you benchmark that using only past data, I think it is by definition the most accurate system. (The data is always a better fit to itself than a theoretical fit is.) Naturally future data is much harder to deal with than past data. But even for future data it's not obvious that ELO (or any other theoretical fit to the odds of winning) will be more accurate than the historical odds.
- deleted 6y ago[deleted]
- BSTRhino 6y agoYes, the best fit for the data is the data itself, it's a tautology. Nothing wrong with Elo's exponential curve, it just can't beat the actual data. You raise a good point in that I could've created a training set and a test set, that probably would be a better validation. But I don't know, I'm not doing science, I'm making a game. On the topic of whether the future matches the past, the predictions were based on a rolling database of the past 100000 matches, which is approximately the number of matches played per 7 days. So my theory is that the data is quite recent and up-to-date and so should match, in general. Of course I never tested this. In the end, I'm not doing science, I'm making a game. If the retention goes up, complaints are down, then I can't keep working on the rating system, there are 1000 other things to do.
- mcnamaratw 6y agoYeah, I'm not giving advice on how you should do it. I was just unsure whether critics here had understood that measured data is probably better than any theoretical fit, even the revered ELO.
- ponker 6y agoWell, it’s like the question of what is better: a restaurant with 4.5 stars on 4 reviews or one with 4.2 stars on 1,500 reviews?
- prionassembly 6y agoIs the point really predicting outcomes? FIDE (chess) Elo is useful because I can compare machines to humans who have never matched each other. Generally speaking the "rating structure" is a lattice where you can, for any two players A and B, tell whether A is a better player than B or the other way around. Elo, Glicko, etc. are embeddings of this lattice on the real line (much like the utility functions of microeconomics are real embeddings of preference lattices).
- defertoreptar 6y ago> Is the point really predicting outcomes? Others have pointed out how there is a psychological aspect of rating systems, and no developer wants to constantly field complaints. That said, I believe the answer is yes. A rating system derives meaningfulness from its predictive power. In other words, people want to know how good they actually are compared to one another.
- Godel_unicode 6y agoI think that for most game players, outside of the top-N group who just want to be at the head of the list, rankings are largely a mechanism to facilitate playing good games, where good is generally defined as close games where both players feel like they could have won. There's an interesting question about how you rate players who use fundamentally different strategies. For instance in RTS games, should you match boom vs blitz players of otherwise equivalent elo? Or should you instead (try to) construct a classifier to determine which type of player someone is, and then have a rank against each other type of player and match them according to that rank?
- ajuc 6y agoAs long as you let players play several games vs each other it's fine. If you met this cannon rushing guy already you'll know what to expect. That's why people use barcode names (llll1l1l1l1l111l) in starcraft :)
- mlyle 6y ago
- naravara 6y agoFor a lot of online games I think matchmaking based on pushing you towards a 50/50 win rate is kind missing the point of games. It gives you fair odds of winning, but it doesn’t necessarily give you even odds of having a fun or competitive game. At high skill levels players are skilled enough to where it might, but with most online multiplayers the overwhelmingly vast majority of players are lacking in basic fundamentals to varying degrees. At that level, ELO based matchmaking mostly just results in one person getting rolled or doing the rolling. They’re not really competitive games in my experience.
- Godel_unicode 6y agoIf two players of similar skill general roll one another, that's a game design problem not a rankings problem.