31 ms·
Riot Games: Artificial Latency for Remote Competitors
- gareth_untether 4y agoReminds me of the cable lengths for black boxes connected to the network in Wall Street. Each cable is the same length regardless of which computer is closer to the access point.
- Misdicorl 4y agoPerhaps apocryphal/silly, but amusing nonetheless. Story goes that this means you want to be in the computer furthest from the interconnect because light travels slightly faster in straight fiber than in coiled fiber.
- colejohnson66 4y agoNot apocryphal; IEX has 38 miles of wire in their building. Tom Scott did a video a few years ago about it: https://www.youtube.com/watch?v=d8BcCLLX4N4 https://www.youtube.com/watch?v=d8BcCLLX4N4
- vmception 4y agoDoes IEX have any liquidity or uptake yet?
- infinityio 4y agoApparently they have 2.6% market share for trading in the US?
- Nition 4y agoI think Misdicorl meant that their story of wanting to be the furthest away is perhaps untrue, not the cables.
- Humdeee 4y agoI work in this space and although true (this is called modal dispersion), there are compensation techniques always used in termination equipment to 'handicap' or mitigate these occurrences. Not perfect, but very, very, very small deltas. Unsure if it's enough to act upon without knowing the length from output to next input.
- jakey_bakey 4y agoThat's neat, I didn't know about this. Very cool.
- jedberg 4y agoI never understood why they didn't use randomized length micro batches to solve this. Instead of processing orders instantly, wait between 200ms and 500ms and then process all orders that came in that window in random order. Then being 5ms closer to the server wouldn't matter.
- axg11 4y agoThat sounds like a complex solution. Sometimes a dumb solution that works good enough is better than a complex solution that _probably_ can't be exploited.
- jedberg 4y agoIt's complex on one end to reduce complexity on the other -- the trading companies wouldn't have to worry about millisecond optimizations if they trading batches were 200ms windows. So the wire lengths wouldn't matter but also not mattering is the processor, memory, software, etc. for the trading companies. Seems like a good tradeoff. And honestly the wire thing probably isn't real. Light moves 30cm in a nanosecond. The wire lengths could be 3 meters different and only make a 10ns difference. Not sure that would matter all that much.
- axg11 4y agoI used to work in HFT. I promise you that companies would still try to exploit randomized batches. There is an advantage to being the very last entrant into a batch (most up to date information). Truly random batches are not trivial to implement and any statistical pattern in the batching could be exploited.
- jerf 4y ago"Truly random batches are not trivial to implement and any statistical pattern in the batching could be exploited." If an HFT manages to exploit a correctly implemented entropy-pool random number generator using AES-256 to extend the stream as needed, they're welcome to pick up a few more bucks as far as I'm concerned. Yes, problems have existed before but I'm sure in this case we can assume high-assurance and careful programming, not some fresh grad being assigned the problem and thinking linear congruential generators sound really cool.
- jannyfer 4y agoThe part I was interested about the most was glossed over. In what situation can this happen? > we realized that there was a calculation error that only manifested in scenarios where the actual ping was significantly lower than the target latency. In this situation the actual latency would be considerably higher than what is displayed on the overlay on the player's screens.
- wswope 4y agoKnowing the classic Riot Games MO: I’d bet money there was a plus instead of a minus that made it through code review, and they’re being intentionally vague to save face.
- bobogei81123 4y agoMy guess is that they did not add in the client delay.
- upwardbound 4y agoExactly right. They added too much latency because they forgot to account for the latency they added in the client. You can see from their architecture diagram (Figure 4) that the latency measurement didn't include the client delay. https://images.contentstack.io/v3/assets/bltad9188aa9a70543a/blt0d8af56e436d6dc3/6283653ad5f9386926e63767/Asset4.PNG https://images.contentstack.io/v3/assets/bltad9188aa9a70543a... The blog post states: "The existing network monitoring system measured the latency at the networking layer as shown by the green arrow."
- jannyfer 4y agoThat kind of error would apply all the time. But the blog post states: "a calculation error that only manifested in scenarios where the actual ping was significantly lower than the target latency".
- upwardbound 4y agoSadly, my guess is that they are transparently lying about that. Since they apply the lag half on the client and half on the server, their lag compensation would have been off by a factor of 2x (or a factor of 100% from a relative perspective) and so they might be just claiming that a 100% error isn't very large when the total lag difference is small. "Yes we were supposed to add 2ms and we added 4ms instead, but at the end we were still only wrong by 2ms (in the other direction) which is not a big deal."
- hsnewman 4y agoDecwars used to "equalize" latency for consoles verses dialups (in the 1970's)...
- NelsonMinar 4y agoWow for real? That's amazing. I'd love to read more about that, it seems quite innovative!
- Kuinox 4y agoSeems something very hard to do in software but very easy to do in hardware.
- bob1029 4y agoI'd argue the complexity of trying to achieve this at all is too much to bear in practice, especially considering the type of customers this would be inflicted upon. Competitive gamers will now also be wondering if the "fake lag" system is bugged, in addition to all of the other problems that could still exist. Some problems cannot be solved with clever tricks.
- infinityio 4y agoif they pulled an IEX and just ran the Busan clients through a few km of fibre to induce delay they could use native pings and verify easily
- Kuinox 4y agoYes, but they probably run on AWS and can't use their own hardware.
- bostonsre 4y agoThink other people have already done the hard part. Wouldn't be too hard to write something that wraps tc and ping. # hping -S -p 1234 somehost.com HPING somehost.com (ens3 1.2.3.4): S set, 40 headers + 0 data bytes len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=0 win=26883 rtt=0.9 ms len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=1 win=26883 rtt=0.8 ms len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=2 win=26883 rtt=0.7 ms # tc qdisc add dev ens3 root netem delay 25ms # hping -S -p 1234 somehost.com HPING somehost.com (ens3 1.2.3.4): S set, 40 headers + 0 data bytes len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=0 win=26883 rtt=25.9 ms len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=1 win=26883 rtt=25.8 ms len=46 ip=1.2.3.4 ttl=64 DF id=0 sport=1234 flags=SA seq=2 win=26883 rtt=25.7 ms
- 4y ago
- dematz 4y agoAs context this was a big international fiasco. Nobody was acting maliciously but fans of RNG and other teams felt wronged. RNG didn't want to attend due to covid restrictions on travel but Riot really wants Chinese viewers in international tournaments, hence all this online latency work. Pros on other teams were lied to by the false latency information which must be almost paranoia inducing to be told it's only 35ms even when it (correctly) feels higher. In the end RNG had to replay their games, which was not as bad as it could have been since they were vs much weaker teams, but still annoying. Anyways, by "scenarios where the actual ping was significantly lower than the target latency" do they mean like if the ping from Busan-server is 10 but they want it to feel like 35? I feel like that's the one scenario you'd care about, making the Busan players have 35 ping was the whole point. Edit - also the other scenario, where the actual ping is significantly higher than the target latency...well in that case your latency service would be some sort of magical time travel box! Adding latency is literally the only thing it can do right? It's not possible to lower it. Maybe the it depends what they mean by "significantly" but idk the target latency was 35ms so like at most they could be 35ms lower than that. I guess they didn't catch it because the bug was in the latency service itself which gave the wrong value to both the in game display and to the old network monitoring system. Idk in hindsight it's easy to say if you're going to introduce this new ping equalizing service it might break how you test ping especially because it hasn't been tested in soloq. But meh a lot easier to see what happened in hindsight after they write it all up.
- vikingerik 4y agoYou can in fact have that magical time-travel approach to get apparent latency lower than the real number. Keep a history of the game state at every tick, so you can roll back to any previous instant. If real latency from one client is 50 ms, and you're targeting 35 ms, you roll back the game state by 15 ms each time input comes in from that client, and re-run the game logic from that point forwards and send the resulting state update to all clients. Those clients will see a bit of jumpiness in the game state, which isn't ideal, but may be the lesser evil compared to enduring more latency. (I have no idea what LoL in particular supports or implements, but that's the general idea.)
- 4y ago
- whateveracct 4y agoA good software lesson here. They built a complex system to tweak tens of ms (at most) of ping to equalize. And they had a bug and it disadvantaged on team. Thus, they have to replay due to unfairness introduced by Riot itself. They could've gone Option 1 - all teams at their natural ping. While it wouldn't be as perfect in theory, it wouldn't have resulted in replaying matches. Plenty of other esports and fighting games play at natural ping and have real results. There is unfairness due to location, but only scrubs blame 5-15ms ping advantage for a loss in a game as strategic as LoL. People play Melee with ping differences bigger than that and it works fine lol. Waaay more technical game shows that there's tolerance. Not to mention that the server option could've had some system to distribute across advantageous servers over the course of a series.
- gosukiwi 4y agoThere is a huge difference between 0 ping and 35 ping though, particularly in pro-level LoL. It is strategic but that ping can be the difference between flashing a skillshot or not, and swinging the whole course of a teamfight. Of course 35 ping is not THAT bad and I agree, while unfair, it wouldnt be the end of the world, it's not even Worlds, just MSI, an invitational event.
- xboxnolifes 4y ago> People play Melee with ping differences bigger than that and it works fine lol. Waaay more technical game shows that there's tolerance. Sure, people play with ping disparity in tons of competitive games, but that's because there's not really an alternative for online play. It's "play with a ping disparity" or "don't play at all". Once there is more widespread reliable alternative, which Riot seems to be working toward in their own game tournaments, maybe we'll see if this tolerance persists in competitive tournaments.
- tester756 4y agoI still don't understand they didn't let anyone proficient at LoL even try using it?
- colinmhayes 4y agoThe difference was .01 seconds. It's not immediately obvious that there is a problem.
- RheingoldRiver 4y agoIt was immediately obvious to competitors and reported on day 1 of competition.
- joebob42 4y agoI suck at lol, the amount of ping it takes for me to be able to tell it's different is only like 30-40 ms. Pro players being able to tell and care at half that seems pretty likely.
- tester756 4y ago>The difference was .01 seconds. It's not immediately obvious that there is a problem. in games you really feel ping differences if you were playing e.g for a year on 5ping and then jumped on 30, then you'll feel that in games like LoL and probably CSGO? idk about the second one.
- colinmhayes 4y agoThey said the ping was 45 instead of 35. I'm not denying that the pros can tell the difference, I'm saying it wouldn't be obvious to whatever tester they had that it was 10ms off.
- Karrot_Kream 4y agoAs others have said, the pros noticed. Having an incorrect calculation and seeing large jitter spikes can probably push the latency into something like 50ms momentarily, enough for pros to lose their "immersion".
- sdwr 4y agoI love the intent and actions behind this post. Finding an unfair bug mid-tournament is serious, and it takes courage and integrity on the part of the technical team to -listen to the players -investigate past what your own analytics are saying -find and fix the bug under time pressure -and disclose the (potentially embarrassing!) initial mistake I think the article itself could have been worded better though. It didnt feel very clear. As a reader, I want to know what went wrong, how it was handled, and feel confident that you are on top of the issue. All the stuff about "why we chose our network topography" reads as details at best, justification/excuses at worst.
- bigcat12345678 4y agoYep, look at what Valve is doing to dota 2...
- gverrilla 4y agowould you care to elaborate? I didn't play dota in a few years
- bigcat12345678 4y agoI don't have a close observation either. I have read r/dota2 regularly, and the laziness of Valve just popped fairly often. Things like items promised got delayed and buyers actually forgot them, serious bug in TI (https://www.reddit.com/r/DotA2/comments/9aexks/a_morphling_bug_that_possibly_changed_the_result/ https://www.reddit.com/r/DotA2/comments/9aexks/a_morphling_b...) etc.
- DantesKite 4y agoThe technical team acted honorably, but the upper stewardship at Riot games behaved horrendously. They sacrificed the competitive integrity of their sport because they didn't want to exclude RNG from the tournament, because they wanted the audience of the Chinese market. This article goes into more detail, giving a larger context for why this situation shouldn't have never happened in the first place. https://www.invenglobal.com/articles/17207/montecristo-its-embarrassing-the-level-of-competitive-integrity-thats-been-sacrificed-for-msi-2022 https://www.invenglobal.com/articles/17207/montecristo-its-e...
- NelsonMinar 4y agoSome data: 8 years ago someone found going from 35ms to 0ms latency meant you were likely to win games 1-2% more of the time: https://www.reddit.com/r/dataisbeautiful/comments/1t23a0/latency_lag_vs_win_rate_in_league_of_legends_oc/ https://www.reddit.com/r/dataisbeautiful/comments/1t23a0/lat... And this is a little different, but Riot found that playing on ethernet instead of wifi made you about 1% more likely to win games. https://web.archive.org/web/20160814131032/http://na.leagueoflegends.com/en/page/ethernet-vs-wifi-ping-packets-playing-better https://web.archive.org/web/20160814131032/http://na.leagueo...
- Karrot_Kream 4y agoI'm going to guess this effect is lower because it's being marginalized across all player skill levels (ELO). In Riot's post you can see that the higher tier the player is at, the more likely they are to use Ethernet. If you conditioned on ELO I suspect you'll find a much larger effect.
- frozenlettuce 4y agoWarcraft 3's (pre-reforged) Battle.net had a 250ms delay for all players (when it was launched dialup connections were still common)
- someweirdperson 4y agoArtificial or natural delays, I am somewhat confused that one end measuring/calculating ping can end up with a wrong value. Transmit at local time x, receive at local time y, subtract. It only needs a working local clock. Second surprising thing is artifical latency added on the client. Anything on the client I would avoid, if possible. Third, this all looks like magically connecting two ends. But it is packets flowing from one to the other, and the other to the one. Latency could be added to each independently, even on one end, by artificially buffering to-be-send and/or received packets. What are the implications for game-play with an asymetric latency, e.g. theoretical 0 delay to receive vs. long delay to transmit, or vice versa?
- chabons 4y agoI think the symmetric delay here makes sense, since you're trying to simulate the latency of a (likely) symmetric connection between another client and the server. If you only add latency on the transmission, then 2 actions taken at the same time by players with different amounts of physical latency will be processed at different times. If we assume the game clocks on each client and the server are in sync, then this creates a competitive advantage for the player with lower physical latency. For instance, if we have a timing battle (ie: Zhonya's hourglass), where both players know that at time 'T' they must each take an action, and the first to do so comes out on top, then the player with higher physical latency is at a disadvantage.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- bcrosby95 4y agoYou can accomplish this by putting it all on the client, all on the server, or a mix of both (what they chose). For all on the server, you could add latency right after receiving data from the client, but before game code processes it, and before sending data to the client. That said, there likely were technical considerations involved in their decision to do a mix of both.
- mleonhard 4y agoThey used a complex and buggy solution when a simple solution exists: Use the `tc` traffic control [0] program to configure the Linux kernel to add a fixed amount of latency to traffic from the local players [1]. If the game server does not run on Linux, they could put a Linux bridge/router in front of it. [0] https://en.wikipedia.org/wiki/Tc_%28Linux%29 https://en.wikipedia.org/wiki/Tc_%28Linux%29 [1] https://serverfault.com/a/841865 https://serverfault.com/a/841865
- connor4312 4y agoAs I understand the blog post, the difficult (and buggy) part is not the addition of latency, but calculating how much latency to add and where to add it. I'm not sure how tc would help much here, and actually don't see anything to indicate they weren't using tc already.
- spmurrayzzz 4y ago`tc` would give you non-deterministic microbursts that would be hard to quantify as to whether they impacted competitive integrity. This is an incredibly complex problem if they actually care about millisecond level precision in small duty cycles.
- mleonhard 4y agoAre you saying that Linux's packet delay function is imprecise? Can you link to more information about this?
- zorgmonkey 4y agoThe problem is that you can't just add a fixed amount of latency to hit a goal of 35ms, you need some sort of algorithm that adaptively changes the amount of added latency to account for occasional spikes in network latency (this is a game that takes >30min so those spikes are guaranteed to happen over the course of a game)
- mleonhard 4y ago
- joshuak 4y agoDistributed consensus is hard. Very hard. Anyone dismissing the problem as “should never have hopped,” hasn’t a clue as to how hard and error prone reliable distributed consensus is. 35 ms lag in overall responsiveness is an eternity in competitive game play. There are videos floating around that show drawing on a tablet surface with various input latencies (perhaps someone has a link, I can’t find them at the moment). 35ms latency is very noticeable to anyone never mind professional competitors. Because of the tricks game designers must pull off in lieu of proper distributed consensus which has hard requirements bound by the laws of physics, it is completely likely there are lots of bugs in the system. I think Riot did the best anyone could reasonably have expected of them, and the write up is particularly informative and helpful.
- joebob42 4y agoThere's no consensus involved in the bug.
- joshuak 4y agoSure there is. The need to introducing a unified delay across all game clients in the first place, reliable measurement and comparison of latency data, and the effects of latency, artificial or otherwise, on game state which is not strongly consistent. There could be many other areas in which this bug intersects with distributed consensus in the actual implementation as well. In fact neglecting the impact of distributed consensus is one of the biggest challenges to mitigating it.
- namibj 4y agoI think you meant https://youtu.be/vOvQCPLkPt4 https://youtu.be/vOvQCPLkPt4 . It was indeed not too easy to find, due to SEO tactics polluting YT search with gaming-oriented stuff from recent times (where one may actually reach the frame rates necessary for such low latencies in a frame-by-frame display technology).
- joshuak 4y agoThanks! That is the one I was thinking of. Notice how even at 10ms of latency rubberbanding is easily visible.
- upwardbound 4y agoIt looks like the error they made was this: They added too much latency because they forgot to account for the latency they added in the client. You can see from their architecture diagram (Figure 4) that the latency measurement didn't include the client delay. https://images.contentstack.io/v3/assets/bltad9188aa9a70543a/blt0d8af56e436d6dc3/6283653ad5f9386926e63767/Asset4.PNG https://images.contentstack.io/v3/assets/bltad9188aa9a70543a... The blog post states: "The existing network monitoring system measured the latency at the networking layer as shown by the green arrow."
- chabons 4y agoI read that section as comparing what they logged before vs what the player experiences: the time taken for an input to be registered and it's effect communicated to the user. Maybe I'm wrong, but I don't think the intention was a precise description of the error, but rather the hazard of not having logs which reflect what the user experiences.
- upwardbound 4y agoThey stated that they use the same calculation for the logs as for other purposes such as latency compensation. From near the top of the blog post: "The reason we did not find it sooner is that the cause of the issue was a code bug that miscalculated latency, which meant that the values in our logs were also wrong." And from later in the post: "Our logs were not displaying the issue because the calculation was wrong. It explained why the latency was worse in the venue than on the internet servers"
- chabons 4y agoRight, but my point was that I think you're reading into the diagram beyond it's intended purpose. I'm not sure how you can tell that they're logging only the server-side delay, and not the client-side delay, from a systems diagram which shows each of those as a box connected by a bi-directional arrow. For instance, it's clear to me that they don't include the latency included by the game engine in that metric, but not whether they include the server-added latency or if they're aware of it and subtract it off.
- Lapsa 4y agoI'm into music production. Try to process some sound through the compressor and adjust its attack - my ears quite easily can discern difference as small as 2ms. Another exercise - try to sing in a group over the internet (you gonna rage quit in five minutes guaranteed).