5 ms·
It's weird it took them so long to disable streaming. One of the first things you do in this case is roll back the last software and config updates, even innoce
by regnull 5y ago
It's weird it took them so long to disable streaming. One of the first things you do in this case is roll back the last software and config updates, even innocent looking ones.
- rkuykendall-com 5y agoHindsight is 20-20 I shouldn't have drank that many Hindsight is 20-20 Stop. – Little elevators are far too small for me So I ride the big ones It's not so fun unless you're OCD And you like buttons
- yashap 5y agoThat’s what stood out to me too. Although they’d been slowly rolling it out for awhile, their last major rollout was quite close to the outage start: > Several months ago, we enabled a new Consul streaming feature on a subset of our services. This feature, designed to lower the CPU usage and network bandwidth of the Consul cluster, worked as expected, so over the next few months we incrementally enabled the feature on more of our backend services. On October 27th at 14:00, one day before the outage, we enabled this feature on a backend service that is responsible for traffic routing. As part of this rollout, in order to prepare for the increased traffic we typically see at the end of the year, we also increased the number of nodes supporting traffic routing by 50% Consul was clearly the culprit early on, and you just made a significant Consul-related infrastructure change, you’d think rolling that back would be one of the first things you’d try. One of the absolute first steps in any outage is “is there any recent change we could possibly see causing this? If so, try rolling it back.” They’ve obviously got a lot of strong engineers there, and it’s easy to critique from the outside, but this certainly struck me as odd. Sounds like they never even tried “let’s try rolling back Consul-related changes”, it was more that, 50+ hours into a full outage, they’d done some deep profiling, and discovered the steaming issue. But IMO root cause analysis is for later, “resolve ASAP” is the first response, and that often involves rollbacks. I wonder if this actually hindered their response: > Roblox Engineering and technical staff from HashiCorp combined efforts to return Roblox to service. We want to acknowledge the HashiCorp team, who brought on board incredible resources and worked with us tirelessly until the issues were resolved. i.e. earlier on, were there HashiCorp peeps saying “naw, we tested streaming very thoroughly, can’t be that”?
- otterley 5y agoWhen you're at Roblox's scale, it is often difficult to know in advance whether you will have a lower MTTR by rolling back or fixing forward. If it takes you longer to resolve a problem by rolling back a significant change than by tweaking a configuration file, then rolling back is not the best action to take. Also, multiple changes may have confounded the analysis. Adjusting the Consul configuration may have been one of many changes that happened in the recent past, and certainly changes in client load could have been a possible culprit.
- yashap 5y agoSome changes are extremely hard to rollback, but this doesn’t sound like one of them. From their report, sounds like the rollback process involved simply making a config change to disable the streaming feature, it took a bit to rollout to all nodes, and then Consul performance almost immediately returned to normal. Blind rollbacks are one thing, but they identified Consul as the issue early on, and clearly made a significant Consul config change shortly before the outage started, that was also clearly quite reversible. Not even trying to roll that back is quite strange to me - that’s gotta be something you try within the first hour of the outage, nevermind the first 50 hours.
- deleted 5y ago[deleted]
- mypalmike 5y agoIn most cases, if you've planned your deployment well (meaning in part that you've specified the rollback steps for your deployment) it's almost impossible to imagine rollback being slower than any other approach. When I worked at Amazon, oncalls within our large team initially had leeway over whether to roll backwards or try to fix problems in situ ("roll forward"). Eventually, the amount of time wasted trying to fix things, and new problems introduced by this ad hoc approach, led to a general policy of always rolling back if there were problems (I think VP approval became required for post-deploy fixes that weren't just rolling back). In this case, though, the deployment happened ages (a whole day!) before the problems erupted. The rollback steps wouldn't necessarily be valid (to your "multiple confounding changes" point). So there was no avoiding at least some time spent analyzing and strategizing before deciding to roll back.
- Twirrim 5y agoThe post indicates they'd been rolling it out for months, and indicate the feature went live "several months ago". With the behaviour matching other types of degradation (hardware), it's entirely reasonable that it could have taken quite a while to recognise that software and configurations that have proven stable for several months, that is still there working, wasn't quite so stable as it seemed.
- fullsend 5y agoHonestly I would guess part of it is that streaming is supposed to be a performance increase. So during a performance related outage, it might be easy to overlook. Am I really going to turn off a feature that I think is actually helping the problem?
- sidlls 5y agoIf that feature was the one most recently deployed or updated? Yes, if possible. That could be a big if, though right? Maybe rolling back such a change isn't trivial, or imposes other costs to returning to service that are more expensive than simply working through the problem.
- captaincaveman 5y agoWell its a feature that affects the specific area where your having issues (performance), so yes would be the right thing to start with.
- brobinson 5y agoThe htop screenshot was an immediate, appropriately-colored red flag for me: that much red (kernel time) on the CPU utilization bars for a system running etcd/consul is not right in my experience.
- atmosx 5y agoSome comments: - The write up is amazing. There is a great level of detail. - When they had the first indication of a problem, instead of looking if the problem was the hardware (disk I/O, etc.) the team went full cattle/cloud: bring down the node, launch a new one. Apparently that cost them a few hours. We would probably have done the same but I wonder if there's a lesson there. - The obvious thing to do was revert configs. It is very strange that it took so long to revert. After being down for hours and having no idea what gives, it's the reasonable thing to try. - The problem was consul. But consul is a key component and Roblox seem to be running a fairly large infrastructure. The company's valuation is sky-high, I assume the infra team is quite large. Consul is an open source project. Wouldn't make sense instead of relying on hashicorp so heavily to bring-in or train ppl around consul internals at this point? (maybe not possible/feasible/optimal, just wondering) Would be a nice touch to check if bbolt has the bug and possibly push a fix. That said, the post-mortem is state-of-art. Way better than anything we've seen from much much bigger companies.
- geoelkh 5y agoThe post mortem is really well written but I had the same thoughts. They upgraded the machines hardware before rolling back the latest config updates.