20 ms·
Roblox October Outage Postmortem
- zomglings 5y agoThis is a great post-mortem - thank you to the Roblox engineering team for being this transparent about the issue and the process you took to fix it. It couldn't have been easy and it sounds like it was a beast to track down (under pressure no less). gg
- erwincoumans 5y ago>> We are working to move to multiple availability zones and data centers. Surprised it was a single availability zone, without redundancy. Having multiple fully independent zones seems more reliable and failsafe.
- kreeben 5y ago>> Having multiple fully independent zones seems more reliable I don't think these independent zones exist. See AWS's recent outages, where east cripples west and vice versa.
- count 5y agoThat's not how they work. They exist, and work extremely well within their defined engineering / design goals. It's much more nuanced than 'everything works independently'.
- kreeben 5y agoIf the design goal of these zones is that they should be independent of each other then, no, they do not work extremely well.
- Karrot_Kream 5y agoAvailability Zones aren't the same thing as regions. AWS regions have multiple Availability Zones. Independent availability zones publishes lower reliability SLAs so you need to load balance across multiple independent availability zones in a region to reach higher reliability. Per AZ SLAs are discussed in more detail here [1] (N.B. I find HN commentary on AWS outages pretty depressing because it becomes pretty obvious that folks don't understand cloud networking concepts at all.) [1]: https://aws.amazon.com/compute/sla/ https://aws.amazon.com/compute/sla/
- kreeben 5y ago>> you need to load balance across multiple independent availability zones The only problem with that is, there are no independent availability zones. What we do have, though, is an architecture where errors propagate cross-zone until they can't propagate any further, because services can't take any more requests, because they froze, because they weren't designed for a split brain scenario, and then, half the internet goes down.
- outworlder 5y ago> The only problem with that is, there are no independent availability zones. There are - they can be as independent as you need them to be. Errors won't necessarily propagate cross-zone. If they do, someone either screwed up, or they made a trade-off. Screwing up is easy, so you need to do chaos testing to make sure your system will survive as intended.
- kreeben 5y agoI'm not talking about my global app. I'm talking about the system I deploy to, the actual plumbing, and how a huge turd in a western toilet causes east's sewerage system to over-flow.
- mlyle 5y ago> (N.B. I find HN commentary on AWS outages pretty depressing because it becomes pretty obvious that folks don't understand cloud networking concepts at all.) What he said was perfectly cogent. Outages in us-east-1 AZ us-east-1a have caused outages in us-west-1a, which is a different region and a different AZ. Or, to put it in the terms of reliability engineering: even though these are abstracted as independent systems, in reality there are common-mode failures that can cause outages to propagate. So, if you span multiple availability zones, you are not spared from events that will impact all of them.
- Karrot_Kream 5y ago> Or, to put it in the terms of reliability engineering: even though these are abstracted as independent systems, in reality there are common-mode failures that can cause outages to propagate. It's up to the _user_ of AWS to design around this level of reliability. This isn't any different than not using AWS. I can run my web business on the super cheap by running it out of my house. Of course, then my site's availability is based around the uptime of my residential internet connection, my residential power, my own ability to keep my server plugged into power, and general reliability of my server's components. I can try to make things more reliable by putting it into a DC, but if a backhoe takes out the fiber to that DC, then the DC will become unavailable. It's up to the _user_ to architect their services to be reliable. AWS isn't magic reliability sauce you sprinkle on your web apps to make them stay up for longer. AWS clearly states in their SLA pages what their EC2 instance SLAs are in a given AZ; it's 99.5% availability for a given EC2 instance in a given region and AZ. This is roughly ~1.82 days, or ~ 43.8 hours, of downtime in a year. If you add a SPOF around a single EC2 instance in a given AZ then your system has a 99.5% availability SLA. Remember the cloud is all about leveraging large amounts commodity hardware instead of leveraging large, high-reliability mainframe style design. This isn't a secret. It's openly called out, like in Nishtala et al's "Scaling Memcache at Facebook" [1] from 2013! The background of all of this is that it costs money, in terms of knowledgable engineers (not like the kinds in this comment thread who are conflating availability zones and regions) who understand these issues. Most companies don't care; they're okay with being down for a couple days a year. But if you want to design high reliability architectures, there are plenty of senior engineers willing to help, _if_ you're willing to pay their salaries. If you want to come up with a lower cognitive overhead cloud solution for high reliability services that's economical for companies, be my guest. I think we'd all welcome innovation in this space. [1]: https://www.usenix.org/system/files/conference/nsdi13/nsdi13-final170_update.pdf https://www.usenix.org/system/files/conference/nsdi13/nsdi13...
- Bluecobra 5y ago> I don't think these independent zones exist. Wouldn't it be possible to create fully independent zones with multiple cloud providers, like AWS, GCP, Azure? This is assuming that your workloads don't rely on proprietary services from a given provider.
- dpifke 5y agoYes, and would also protect you from administrative outages like, "AWS shut off our account because we missed the email about our credit card expiring." (But wouldn't protect you from software/configuration issues if you're running the same stack in every zone.)
- vorpalhex 5y agoI'm more impressed that it hasn't been an issue until now.
- bob1029 5y ago> Having multiple fully independent zones seems more reliable and failsafe. This also introduces new modes of failure which did not exist before. There are no silver bullets for this problem.
- rhizome 5y agoThere are no silver bullets to any problem, but there are other ways of implementing services and architecture that can sidestep these things.
- hedwall 5y agoA guess would be that game servers are distributed across the globe but backend services l are in one place. A common pattern in game companies.
- foobarian 5y ago> Surprised it was a single availability zone, without redundancy. Having multiple fully independent zones seems more reliable and failsafe. It's also a lot more expensive. Probably order of magnitude more expensive than the cost of a 1 day outage
- outworlder 5y ago> It's also a lot more expensive. Probably order of magnitude more expensive than the cost of a 1 day outage Not sure I agree. Yes, network costs are higher, but your overall costs may not be depending on how you architect. Independent services across AZs? Sure. You'll have multiples of your current costs. Deploying your clusters spanning AZs? Not that much - you'll pay for AZ traffic though.
- adrr 5y agoIt is when you run your own date centers and have to shell out a large capital outlays to spin up a new datacenter.
- Symbiote 5y agoThe usual way this works (and I assume this is the case for Roblox) is not by constructing buildings, but by renting space in someone else's datacentre. Pretty much every city worldwide has at least one place providing power, cooling, racks and (optionally) network. You rent space for one or more servers, or you rent racks, or parts of a floor, or whole floors. You buy your own servers, and either install them yourself, or pay the datacentre staff to install them.
- Hamuko 5y agoHow expensive? Remember that the Roblox Corporation does about a billion dollars in revenue per year and takes about 50% of all revenue developers generate on their platform.
- dev_by_day 5y ago
- abarringer 5y agoWas on a call with a bank VP that had moved to AWS. Asked how it was going. Said it was going great after six months but just learning about availability zones so they were going to have to rework a bunch of things. Astonishing how our important infrastructure is moved to AWS with zero knowledge of how AWS works.
- maxclark 5y agoNo surprised at all. Multi AZ is a PITA. You'd be surprised how many 7fig+/month infra is single region/az
- mhitza 5y agoFor example parts of AWS itself. us-east-1 having issues? Looks like aws console all over the world have issues. You constantly hear about multi zone, region, cloud. But in practice when things break you hear all these stories of them running in a single region+zone
- mbesto 5y agoThere have been multiple discussions on HN about cloud vs not cloud and there are endless amount of opinions of "cloud is a waste blah blah". This is exactly one of the reasons people go cloud. Introducing an additional AZ is a click of a button and some relatively trivial infrastructure as code scripting, even at this scale. Running your own data center and AZ on the other hand requires a very tight relationship with your data center provider at global scale. For a platform like Roblox where downtime equals money loss (i.e. every hour of the day people make purchases), then there is a real tangible benefit to using something like AWS. 72 hours downtime is A LOT, and we're talking potentially millions of dollars of real value lost and millions of potential in brand value lost. I'm not saying definitively they would save money (in this case profit impact) by going to AWS, but there is definitely a story to be had here.
- treis 5y agoBut it wasn't a hardware issue. It was a software one and that would have crossed AZ boundaries.
- mbesto 5y agoSo then why does the post mortem suggest setting up multi-az to address the problems they encountered?
- treis 5y agoI took that to mean sharding Roblox instead of spanning it across data center AZs.
- mbesto 5y agoFTA: > Running all Roblox backend services on one Consul cluster left us exposed to an outage of this nature. We have already built out the servers and networking for an additional, geographically distinct data center that will host our backend services. We have efforts underway to move to multiple availability zones within these data centers; we have made major modifications to our engineering roadmap and our staffing plans in order to accelerate these efforts. If they were in AWS they could have used Consul across multi-AZs and done changes in a roll out fashion.
- statguy 5y agoSo the outage lasted 3 days and the postmortem took 3 months!
- encryptluks2 5y ago
- Operyl 5y agoThey just got out of their busiest time of year, and taking the time to write an accurate post mortem with data gleamed afterwards seems sensible to me.
- koshergweilo 5y agoRead the article " It has been 2.5 months since the outage. What have we been up to? We used this time to learn as much as we could from the outage, to adjust engineering priorities based on what we learned, and to aggressively harden our systems. One of our Roblox values is Respect The Community, and while we could have issued a post sooner to explain what happened, we felt we owed it to you, our community, to make significant progress on improving the reliability of our systems before publishing." They wanted to make sure everything was fixed before publishing
- conorh 5y agoExcellent write up. Reading a thorough, detailed and open postmortem like this makes me respect the company. They may have issues but it sounds like the type of company that (hopefully) does not blame, has open processes, and looks to improve - the type of company I'd want to work for!
- digitalengineer 5y ago
- micromacrofoot 5y agoYeah as long as Roblox is exploiting children they're just flat-out not respectable. This video is a good look at a phenomenon most people are unaware of.
- deleted 5y ago[deleted]
- charcircuit 5y agoPlayers of your game creating content for it is not exploitation. It's just how it works in the gaming world. When I was a kid I spent time creating a minecraft mod that hundreds of people used. Did Mojang or anyone else ever pay me? No. I did it because I wanted to.
- jawngee 5y agoMojang was likely not selling you on making a mod with promises of making money though. Roblox did that, maybe they still do it.
- digitalengineer 5y agoPlease review the video. The problem is not ‘players creating content’.
- micromacrofoot 5y ago
- kjw 5y agoI would not have guessed Roblox was on-prem with such little redundancy. Later in the post, they address the obvious “why not public cloud question”? They argue that running their own hardware gives them advantages to cost and performance. But those seem irrelevant if usage and revenue go to zero when you can’t keep a service up. It will be interesting to see how well this architecural decision ages if they keep scaling to their ambitions. I wonder about their ability to recruit the level of talent required to run a service at this scale.
- nomel 5y ago> But those seem irrelevant if usage and revenue go to zero when you can’t keep a service up You're assuming the average profits lost are more than the average cost of doing things differently, which, according to their statement, is not the case.
- otterley 5y agoSince the issue's root cause was a pathological database software issue, Roblox would have suffered the same issue in the public cloud. (I am assuming for this analysis that their software stack would be identical.) Perhaps they would have been better off with other distributed databases than Consul (e.g., DynamoDB), but at their scale, that's not guaranteed, either. Different choices present different potential difficulties. Playing "what-if" thought experiments is fun, but when the rubber hits the road, you often find that things that are stable for 99.99%+ of load patterns encounter previously unforeseen problems once you get into that far-right-hand side of the scale. And it's not like we've completely mastered squeezing performance out of huge CPU core counts on NUMA architectures while avoiding bottlenecking on critical sections in software. This shit is hard, man.
- baskethead 5y agoThis is not true, if they handled the rollout properly. Companies like Uber have two entirely different data centers and during outages they failover you either datacenter. Everything is duplicated which is potentially wasteful but ensures complete redundancy and it’s an insurance policy. If you rollout, you rollout to each datacenter separately. So in this case rolling out in one complete datacenter and waiting a day for their Consul streaming changes probably would have caught it.
- stuff4ben 5y agoSounds like they need to switch to Kubernetes? I kid of course. One of the best post-mortems I've seen. I'm sure there are K8s horror stories out there of etcd giving up the ghost in a similar fashion.
- schoolornot 5y agoThe one thing you can say about Nomad is that's generally incredibly scalable compared to Kubernetes. At 1000+ nodes over multiple datacenters, things in Kube seem to break down.
- tapoxi 5y agoDo they still? GKE supports 15,000 nodes per cluster.
- spydum 5y agoyou joke, but it's precisely this: >Critical monitoring systems that would have provided better visibility into the cause of the outage relied on affected systems, such as Consul. This combination severely hampered the triage process. which gives me goosebumps whenever I hear people proselytizing everything run on Kubernetes. At some point, it makes good sense to keep capabilities isolated from each other, especially when those functions are key to keeping the lights on. Mapping out system dependencies (either systems, software components, etc) is really the soft underbelly of most tech stacks.
- YATA0 5y ago>Sounds like they need to switch to Kubernetes? Hah! Good one!
- ctvo 5y agoIt's a spicy read. Really could have happened to anyone. All very reasonable assumptions and steps taken. You could argue they could have more thoroughly load tested Consul, but doubtful any of us would have done more due diligence than they did with the slow rollout of streaming support. (Ignoring the points around observability dependencies on the system that went down causing the failure to be extended)
- yashap 5y agoThe main mistake IMO is that, the day before the outage, they made a significant Consul-related infra change. Then they have this massive outage, where Consul is clearly the root cause, but nobody ever tries rolling that recent change back? That’s weird. I went into more detail here: https://news.ycombinator.com/item?id=30015826 https://news.ycombinator.com/item?id=30015826 The outage occurring could certainly happen to anyone, but it taking 72 hours to resolve seems like a pretty fundamental SRE mistake. It’s also strange that “try rollbacks of changes related to the affected system” isn’t even acknowledged as a learning in their learnings/action items section.
- tptacek 5y agoThat doesn't sound accurate. Wasn't the major change they ended up rolling back Consul streaming, which they'd enabled months before, and had been slowly rolling out?
- twblalock 5y agoRight, but the day before the outage, they enabled streaming for a service that didn't have it turned on. That's a discrete config change, the day before the outage.
- yashap 5y ago> Several months ago, we enabled a new Consul streaming feature on a subset of our services. This feature, designed to lower the CPU usage and network bandwidth of the Consul cluster, worked as expected, so over the next few months we incrementally enabled the feature on more of our backend services. On October 27th at 14:00, one day before the outage, we enabled this feature on a backend service that is responsible for traffic routing. As part of this rollout, in order to prepare for the increased traffic we typically see at the end of the year, we also increased the number of nodes supporting traffic routing by 50% So they rolled out a pretty significant Consul related change the day before their massive Consul outage began. They’d been doing a slow rollout, but ramping it up a bunch is a significant change.
- chainwax 5y agoLove the "Note on Public Cloud", and their stance on owning and operating their own hardware in general. I know there has to be people thinking this could all be avoided/the blame could be passed if they used a public cloud solution. Directly addressing that and doubling down on your philosophies is a badass move, especially after a situation like this.
- Neil44 5y agoIt's interesting, I don't see that being on cloud would have avoided or helped this situation much. They were able to ramp up their hardware very quickly - who knows where they got it that fast - and it actually made the problem worse, so being on cloud and having the ability to do that with keystrokes would not have helped. You could say they might be using a different set of components if they were on cloud which may not have suffered the same issues, but you can play the what if game all day it's not related to pros/cons of public cloud.
- regnull 5y agoIt's weird it took them so long to disable streaming. One of the first things you do in this case is roll back the last software and config updates, even innocent looking ones.
- rkuykendall-com 5y agoHindsight is 20-20 I shouldn't have drank that many Hindsight is 20-20 Stop. – Little elevators are far too small for me So I ride the big ones It's not so fun unless you're OCD And you like buttons
- yashap 5y agoThat’s what stood out to me too. Although they’d been slowly rolling it out for awhile, their last major rollout was quite close to the outage start: > Several months ago, we enabled a new Consul streaming feature on a subset of our services. This feature, designed to lower the CPU usage and network bandwidth of the Consul cluster, worked as expected, so over the next few months we incrementally enabled the feature on more of our backend services. On October 27th at 14:00, one day before the outage, we enabled this feature on a backend service that is responsible for traffic routing. As part of this rollout, in order to prepare for the increased traffic we typically see at the end of the year, we also increased the number of nodes supporting traffic routing by 50% Consul was clearly the culprit early on, and you just made a significant Consul-related infrastructure change, you’d think rolling that back would be one of the first things you’d try. One of the absolute first steps in any outage is “is there any recent change we could possibly see causing this? If so, try rolling it back.” They’ve obviously got a lot of strong engineers there, and it’s easy to critique from the outside, but this certainly struck me as odd. Sounds like they never even tried “let’s try rolling back Consul-related changes”, it was more that, 50+ hours into a full outage, they’d done some deep profiling, and discovered the steaming issue. But IMO root cause analysis is for later, “resolve ASAP” is the first response, and that often involves rollbacks. I wonder if this actually hindered their response: > Roblox Engineering and technical staff from HashiCorp combined efforts to return Roblox to service. We want to acknowledge the HashiCorp team, who brought on board incredible resources and worked with us tirelessly until the issues were resolved. i.e. earlier on, were there HashiCorp peeps saying “naw, we tested streaming very thoroughly, can’t be that”?
- jandrese 5y agoThe BoltDB issue seems like straight up bad design. Needing a freelist is fine, needing to sync the entire freelist to disk after every append is pants on head.
- benbjohnson 5y agoBoltDB author here. Yes, it is a bad design. The project was never intended to go to production but rather it was a port of LMDB so I could understand the internals. I simplified the freelist handling since it was a toy project. At Shopify, we had some serious issues at the time (~2014) with either LMDB or the Go driver that we couldn't resolve after several months so we swapped out for Bolt. And alas, my poor design stuck around. LMDB uses a regular bucket for the freelist whereas Bolt simply saved the list as an array. It simplified the logic quite a bit and generally didn't cause a problem for most use cases. It only became an issue when someone wrote a ton of data and then deleted it and never used it again. Roblox reported having 4GB of free pages which translated into a giant array of 4-byte page numbers.
- otterley 5y agoI, for one, appreciate you owning this. It takes humility and strength of character to admit one's errors. And Heaven knows we all make them, large and small.
- klabb3 5y agoI also appreciate the honesty, but I don't see the error in the author, quite the opposite. Afaiu, Bolt is a personal OSS project, github repo is archived with last commit 4 years ago, and the first thing you see in the readme is the "author no longer has time nor energy to continue". Commercial cash cows like Roblox (a) shouldn't expect free labor and (b) should be wise enough to recognize tech debt or immaturity in their dependencies. Heck, even as a solo dev I review every direct dependency I take on, at least to a minimal level. I can't speak to the incident response as I'm not an sre, but as a dev this screams of fragile "ship fast" culture, despite all the back patting in the post. I'm all for blameless postmortems, but a culture of rigor is a collective property worthy of attention and criticism.
- ryanworl 5y agoIt seems that Consul does not have the ability to use the newer hashmap implementation of freelist that Alibaba implemented for etcd. I cannot find any reference to setting this option in Consul's configuration. Unfortunate, given it has been around for a while. https://www.alibabacloud.com/blog/594750 https://www.alibabacloud.com/blog/594750
- throwdbaaway 5y agoI think they just made the switch to the fork that does contain the freelist improvement in https://github.com/hashicorp/consul/pull/11720 https://github.com/hashicorp/consul/pull/11720 Took a major incident to swallow your pride? (consul, powered by go.etcd.io/bbolt)
- ryanworl 5y agoIs this option enabled by default? I don't this it is and I don't think they actually set it manually anywhere. EDIT: I think we're talking about two different options. I meant the ability to leave sync turned on but change the data structure.
- throwdbaaway 5y agoThis is the PR for the freelist improvement from that alibaba article: https://github.com/etcd-io/bbolt/pull/141 https://github.com/etcd-io/bbolt/pull/141 Just to be clear, we are talking about this item from the post-mortem right? > We are working closely with HashiCorp to deploy a new version of Consul that replaces BoltDB with a successor called bbolt that does not have the same issue with unbounded freelist growth. EDIT: I see what you mean. The freelist improvement has to be enabled by setting the `FreelistType` config to "hashmap" (default is "array"). Indeed it doesn't look like consul has done that...
- ryanworl 5y agoI think we're still talking about different things, but that is a good move on their part regardless. :) I mean the optional called `FreelistType` has a new option called `FreelistMapType` and the default is `FreelistArrayType`. There is no option in Consul from what I can tell to configure that option. They did have to upgrade from the old boltdb code to the etcd boltdb code to do this though.
- johnmarcus 5y agoaaaalllllllll the way down at the bottom is this gem: >Some core Roblox services are using Consul’s KV store directly as a convenient place to store data, even though we have other storage systems that are likely more appropriate. Yeah, don't use consul as redis, they are not the same.
- stuff4ben 5y agoBut you can... which is what some engineers were thinking. In my experience they do this because: A) they're afraid to ask for permission and would rather ask for forgiveness B) management refused to provision extra infra to support the engineers need, but they needed to do this "one thing" anyways C) security was lax and permissions were wide open so people just decided to take advantage of it to test a thing that then became a feature and so they kept it but "put it on the backlog" to refactor to something better later
- aprdm 5y agoYes, this and having such a big consul cluster where the recommendation is to have more smaller clusters. That said, could've happened to anyone and it was a great write up.
- AaronFriel 5y agoThis outage has it all, distributed systems, non-uniform memory access contention (aka "you wanted scale up? how about instead we make your CPU a distributed system that you have to reason about?"), a defect in a log-structured merge tree based data store, malfunctioning heartbeats affecting scheduling, wow wow wow. Big props to the on-calls during this.
- tacLog 5y ago> Big props to the on-calls during this. Kind of curious about this. I know this is probably company specific but how do outages get handled at large orgs? Would the on-calls have been called in first then called in the rest of the relevant team? Is their a leadership structure that takes command of the incident to make big coordinated decisions to manage the risk of different approaches? Would this have represented crunch time to all the relevant people or would this be a core team with other people helping as needed?
- sciurus 5y agoApproaches vary company-to-company, but https://response.pagerduty.com/ https://response.pagerduty.com/ is a good resource for understanding how it often looks.
- WaxProlix 5y agoOncalls get paged first and then escalate. As they assess impact to other teams and orgs, they usually post their tickets to a shared space. Once multiple team/org impact is determined, leadership and relevant ops groups (networking, eg) get pulled in to a call. A single ticket gets designated the Master Ticket for the Event, and oncalls dump diagnostic info there. Root cause is found (hopefully), affected teams work to mitigate while RC team rushes to fix. The largest of these calls I've seen was well into the hundreds of sw engineers, managers, network engineers, etc.
- YZF 5y agoSo 1-3 people actually figure it out while everyone else gets in the way? There's no way hundreds of engineers, managers, network engineers etc. can get anything actually done as a group, right?
- samkone 5y ago
- dang 5y agoCould you please stop posting unsubstantive comments to HN? You've done it quite a bit, unfortunately. We're trying for a different quality of discussion here. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- wizwit999 5y ago> On October 27th at 14:00, one day before the outage, we enabled this feature on a backend service that is responsible for traffic routing. As part of this rollout, in order to prepare for the increased traffic we typically see at the end of the year, we also increased the number of nodes supporting traffic routing by 50%. Seems like the smoking gun, this should have been identified and rolled back much earlier.
- yuliyp 5y agoIt's obvious when it's pointed out in an article like this. It's less clear when it's one of many changes that could have been happening in a day, and it was an operation that was considered "safe" given that it had been done multiple times for other services in the preceding months.
- Karrot_Kream 5y agoIf reading a postmortem makes the smoking gun obvious, then the postmortem is doing its job. Don't mistake the amount of investigation that goes into a postmortem for the available information and mental headspace during an outage.
- wizwit999 5y agoI've been in my fair share of incidents so I'm aware of how they work. But they knew it was an issue related to Consul within hours. It shouldn't take more than two days before they check for recent deployments made to Consul.
- orangepenguin 5y agoIn my experience, there's typically more than one "smoking gun". The problem isn't finding one, it's eliminating all of the "smoking guns" that aren't actually related to the outage. If I worked at an organization with many teams deploying updates multiple times per day and several same day events seemed related, I would probably also put less weight on a gradual, months-long deployment that had completed a day prior.
- Twirrim 5y ago"We enjoyed seeing some of our most dedicated players figure out our DNS steering scheme and start exchanging this information on Twitter so that they could get “early” access as we brought the service back up." Why do I have a feeling "enjoyed" wasn't really enjoyed so much as "WTF", followed by "oh shit..." at the thought that their main way to balance load may have gone out the window.
- Symbiote 5y agoIt's difficult to know how quickly word could have spread, but I enjoy knowing a few 11 year olds learned something about the Internet in order to play a game an hour early.
- Twirrim 5y agoWith social media etc, I can see it spreading really fast. That would be my bigger fear trying to get a service back up from a very long outage like that.
- deathanatos 5y agoAt their scale, it was probably an insignificant minority. I read that as nothing more than a wink and nod of "we see what you did ;)" ; which I appreciate. Some companies would have a fit and go nuclear on people for that, for no particular reason. As long as it is an insignificant minority, it doesn't matter, and ideally it's teenagers learning how something works on the side, and that helped grow some future hacker (in the HN sense) somewhere.
- DaiPlusPlus 5y ago> Some companies would have a fit and go nuclear on people for that, for no particular reason Sometimes it's even the Missouri state governor doing that too.
- fragmede 5y agoThe intentionally slow bringup is to handle the thundering herd of having the system come back online to 100% at once. If a couple hundred users (small percentage of userbase) here or there are able to jump to queue, it's no real big deal. As far as players figuring out the DNS steering scheme; the company has no responsibility to keep a non-advertised backend up. If it was a problem, disallow new connection to it and remove it from the main pool.
- willcipriano 5y agoI have this little idea I think about called the "status update chain". When I worked in small organizations and we had issues the status update chain looked like this: ceo-->me, as the organizations got larger the chain got longer first it was ceo-->manager-->me then ceo-->director-->manager-->me and so on. I wonder how long the status update chains are at companies like this? How long does at status update take to make it end to end?
- tacLog 5y agoI am sorry, I didn't have enough context to understand what your saying. When you say: status update chain: ceo --> me. What information is flowing from the CEO to you? or is it the other way around?
- willcipriano 5y agoBoth directions, he is asking "What is going on" and I am telling him. As the org gets larger the request to know what is going on passes down the chain and the reply passes back up.
- NordSteve 5y agoIn well designed incident comms systems, the upward comms occurs automatically, not on request. My goal has always been that my execs know what is going on, so that they are never caught short by status queries.
- sjtindell 5y agoUsually there’s a central place where status is being updated and shared by everyone (a Slack channel for example) and everyone in the chain can just read/ping/respond there. Less of a chain.
- tacLog 5y agoThanks, that makes sense. I haven't experienced that myself yet so I wasn't sure.
- 5y ago
- ineedasername 5y ago">circular dependencies in our observability stack" This appears to be why the outage was extended, and was referenced elsewhere too. It's hard to diagnose something when part of the diagnostic tool kit is also malfunctioning.
- phgn 5y agoLike the Facebook outage a few months ago, when their DNS being down prevented them from communicating interally.
- sjtindell 5y agoSuper interesting. A place where ipvs or ebpf rules per-host for the discovery of services seems much more resilient than this heavy reliance on a functional consul service. The team shared a great postmortem here. I know the feeling well of testing something like a full redeploy and seeing no improvement…easy to lose hope at that point. 70+ hours of a full outage, multiple failed attempts to restore, has got to result in a few grey hairs worth of stress. Well done to all the sre, frontline, support engineers, devs, and whoever else rolled up their sleeves and got after it. The lessons learned here could only have been learned in an infra this big.
- tlynchpin 5y agowarning, completely pedantic pet peeve. > Note all dates and time in this blog post are in Pacific Standard Time (PST). But the incident was during PDT. Just use UTC or colloquial "Pacific time" or equiv and never be wrong! My heart goes out to these people. I can imagine how much sustained terror they were feeling, stare hard and harder at your terminals and still nothing makes sense.
- nanis 5y ago> 50th percentile I would normally not call this out, but it is repeated so often in the text that it is jarring. Just call it "median" as it is everywhere else, please. On the other hand, I must commend the author(s) for not using "based off of" :-) Great write-up, otherwise.
- 867-5309 5y ago>for our most performance and latency critical workloads, we have made the choice to build and manage our own infrastructure on-prem I don't understand this logic. are they basically saying that their servers are on average closer to the user than mainstream cloud infra? are they e.g. choosing to have N satellite servers around a city instead of N instances at one cloud provider location in the centre of the city? is it the sparseness of the servers that decreases the latency? or is it more to do with avoiding the herd, i.e. less trafficky routes / beating the queues? it's also unclear whether they use their own hardware on rented rackspace as that could potentially lower costs too
- mike_d 5y agoCloud providers are rarely in cities. Google's biggest region is in the middle of Iowa, Amazon's is in Virginia. If you have a latency sensitive application (like multiplayer games) it makes sense to put a few servers in each of 100 locations rather concentrate them in a half dozen cloud regions. As they point out elsewhere, the cost of infrastructure directly impacts their ability to pay creators on the platform. Doing it yourself will always be cheaper, and they hired the smart people to make it happen.
- InsomniacL 5y ago> it makes sense to put a few servers in each of 100 locations rather concentrate them in a half dozen cloud regions. Large cloud providers have a backbone network with interconnects to many ISPs reducing the amount of Hops a client has to take across the internet. > Doing it yourself will always be cheaper Treating the Cloud as a traditional IAAS Datacenter extension will be more expensive. By utilising PAAS, only using resource that's needed and when it's needed, etc.. is much cheaper.
- snwfog 5y agoIs there any tutorial on how go get a pref report like the one show in this screenshot? https://blog.roblox.com/wp-content/uploads/2021/11/3-perf-report-1536x787.png https://blog.roblox.com/wp-content/uploads/2021/11/3-perf-re...
- TheDong 5y agoYes. It's the default output for "perf report". I recommend reading this: https://www.brendangregg.com/perf.html https://www.brendangregg.com/perf.html However, the short 2-line way to get that output is the following: perf record -F99 -g --pid $(pidof consul) # Wait a few seconds and hit ctrl-c perf report You'll get similar output to what they show if you have consul running with a similar load :)
- qaq 5y agoLove NATS for not having to deal with service discovery at all.
- NightMKoder 5y agoAdmittedly this is armchair architecture talk, but it seems like either consul or Roblox's use of Consul is falling into a CAP-trap: they are using a CP system when what they need is an eventually-consistent AP system. Granted, the use of consul seems heterogenous, but it seems like the main root cause was service discovery. And service discovery loves stale data. Service discovery largely doesn't change that often. Especially in an outage where a lot of things that churn service discovery are disabled (e.g. deploys), returning stale responses should work fine. There's a reason DNS works this way - it prioritizes having any response, even if stale, since most DNS entries don't change that frequently. That said, DNS is not a great service discovery mechanism for other reasons. Not sure if there's an off-the-shelf solution that relies more on fast invalidation rather than distributed consistent stores.
- tptacek 5y agoCan you say more about service discovery "loving stale data"? Loves in the sense of "generates a lot of it; is constantly plagued by it"?
- boulos 5y agoTheir comment implies "are totally fine with stale data". Their argument is that the membership set for a service (especially on-prem) doesn't change all that frequently, and even if it's out of date, it's likely that most of the endpoints are still actually servicing the thing you were looking for. That plus client retries and you're often pretty good.
- tptacek 5y agoMaybe I'm just working on an idiosyncratic version of the service discovery problem, but "stale data" is basically my bête noire. Part of it is that I don't control all my clients, and I can't guarantee they have sane retry logic; what service discovery tells them is the best place to go had better be responsive, or we're effectively having an outage. For us, service discovery is exquisitely sensitive to stale data.
- Quantumhunk 5y agoWhat I learn from this is issue is partly because of not proper use "Go Channels" and open source product "BoltDB"
- CyanLite2 5y agoThat and going all-in on Hashicorp.
- ekimekim 5y agoIMO looking at the root causes here isn't that helpful. Software is complicated and there will always be some unknown bottleneck or bug lurking to knock you over on a bad day. The important lessons here are about: * How their system architecture made them particularly vulnerable to this kind of issue * Their actions to diagnose and attempt to mitigate the issue * The whole later part about effectively cold-starting their entire infrastructure, all while millions of users were banging on their metaphorical door to start using the service again.
- alpb 5y agoI still don't understand how the elevated Consul latency ended up bringing all of the fleet to halt, fail health checks and drop user traffic. I guess use cases calling consul directly (e.g. service discovery) or indirectly (e.g. vault) could not tolerate having stale reads or sticking with what they've read? If anyone can shed a light on this, I appreciate.
- jeffrallen 5y agoTldr: We made a single point of failure, then we made it super reliable, then the stuff it was doing to maintain itself made it slow itself down, then our single point of failure took down our service. Would be interesting to compare this result to the classic paper on Tandem failures: A. Thakur, R. K. Iyer, L. Young and I. Lee, "Analysis of failures in the Tandem NonStop-UX Operating System," Proceedings of Sixth International Symposium on Software Reliability Engineering. ISSRE'95, 1995, pp. 40-50, doi: 10.1109/ISSRE.1995.497642.
- kalev 5y agoSlightly offtopic; “the team decided to replace all the nodes in the Consul cluster with new, more powerful machines”. How do teams usually do this quickly? Is it a combination of Terraform to provision servers and something like Ansible to install and configure software on it?
- stingraycharles 5y agoTotally depends on how “disciplined” the team’s DevOps practices are. In theory it should be as easy as updating a config parameter as you say, but my experience tells me that it’s sometimes not the case. Especially with these kind of fundamental, core services such as Consul provides, it’s not unheard of to have templates with static machine allocations (as opposed to everything in a single auto-scaling group). It’s a bit of a shortcut, but it’s often a bit hairy to implement these services using true auto-scaling. Having said all this, doing these types of migrations when things are already completely broken / on fire makes things a lot easier: you don’t care about downtime. So then it can be as simple as restarting all instances using a new instance type, downtime be damned.
- oars 5y agoIf I wasn't using AWS I would have no idea how to do this.
- captaincaveman 5y agoSounds like they didn't check what had changed first, before starting to fix things with best guesses ... not saying I wouldn't do the same, but arguably lost them a lot of time.
- k8sToGo 5y agoThey were aware of the changes, but as they stated: it seemed to be working fine and, therefore, was ruled out early on as a potential problem.
- k8sToGo 5y agoDoes anyone know what tool this one is? https://blog.roblox.com/wp-content/uploads/2021/11/3-perf-report.png https://blog.roblox.com/wp-content/uploads/2021/11/3-perf-re... Is it really perf?
- doublerabbit 5y ago> /wp-content/uploads/2021/11/3-perf-report.png It's perf.
- ketanhwr 5y agoIt's perf-report[0] which reads the output of a perf data file and displays the profile. [0]: https://man7.org/linux/man-pages/man1/perf-report.1.html https://man7.org/linux/man-pages/man1/perf-report.1.html
- fifticon 5y agoFor me as a roblox user/programmer, the most annoying part of this was, that their desktop development tools refused to run during this outage, because they insist on "phoning home" when you launch them. It is annoying, because the tools actually run perfectly fine on a local desktop, once you are past the "mothership handshake". I spent that week reading roblox dev documentation instead.
- southerntofu 5y agoWow, that's a very bad trend we see emerging in those past years. Did you have any chance to investigate the request at play and whether you could impersonate it via DNS on your local network (or if on the opposite, TLS certificates were stapled into the app)? Also, i'm curious about your experiences with Roblox. I've only heard about it from these HN threads (no i don't know a single person using it) so if you have feedback to share regarding how to program it and how it compares to a modern game engine/editor like Godot, i'm all ears. Also, if you know of a free-software "alternative" to Roblox ; i'm amazed we run proprietary software in the first place, worried when it doesn't run because it can't phone home, but i'm actually ashamed we end up *producing* (eg. developing) content with proprietary tools that these companies can take away from us any minute.
- londons_explore 5y agoI think this outage was made worse by them not being properly in a big cloud provider. In a cloud provider, having a few people working simultaneously on spinning up instances with different potential fixes, running different tests, and then directing all traffic to the first one that works properly is a viable path to a solution. When you have your own hardware, you can really only try one thing at a time.
- KronisLV 5y ago> When you have your own hardware, you can really only try one thing at a time. How so? What would prevent you from hiring 5-10 people for Ops heavy stuff and getting a bit more hardware resources and doing those things in staging environments with load tests and whatnot? I mean, isn't that how you should do things, regardless of where your infra and software is?
- londons_explore 5y agoIf you own your own hardware, for a given service you probably have enough hardware for the production workload, plus maybe 50% more for dev, test, staging, experiments, etc. All those other environments will probably be scaled down versions. Sure, they can be used in an emergency situation, but they can't withstand the full production load, and anyway they're likely on a separate physical hardware and network (usually you want good isolation between production and test environments). If you use AWS, then you probably on average use the same day to day, but in an emergency you can spin up 5 versions of full production scale to test 5 things at once, and just edit a configfile to direct production traffic to any.
- deleted 5y ago[deleted]
- elfchief 5y agoOne thing I don't see mentioned -- why is the write load so high? Can anyone from Roblox say? (I have a specific reason for asking.)
- fasteo 5y ago>>> The scale of our deployment is significant, with over 18,000 servers and 170,000 containers. That's impressive.
- LaserToy 5y agoJust curious, does Roblox push engineers to learn internals of critical software they operate or they lean on vendors. If vendors, it is reckless.