7 ms·
Downtime last Saturday
- gleb 14y agoIt seems that every GutHub downtime I can recall was caused by automated failover.
- rdl 14y agoI wonder how much downtime they've avoided through automated failover.
- gleb 14y agoYuh, I have the same question too :-). Hard to do cost/benefit analysis when you only see the costs.
- rdl 14y agoIn general, automated failover seems to make most small problems non-problems, but turns some small problems into big problems. It probably depends on actual numbers what makes sense for you app. For some systems, I'd take getting rid of small outages -- I'll happily take an increased risk of a projected 15 minute loss of heart function becoming a >60 minute loss of heart function if it also eliminates what would otherwise be a bunch of 5 minute losses of heart function, since even the 5 minute interruptions would be fatal. (Or, for a better example, revolvers vs. semi-autos. A revolver generally is more reliable, but if it goes out of timing, it's basically doomed, whereas a semi-auto can jam or pieces can break, but a monkey can clear, and a trained monkey can fix.)
- kyrra 14y agoFailover is meant to deal with hardware failures, which will tend to work just fine. But if the node you are failing over onto was already has 60% capacity and you add another 60% capacity during the failover, things are going to get worse. The top-level systems probably need to be able to deal with increased latency or timeouts, and properly handle retries and throttling of traffic. If you have some HA failover setup going but your alternate is already being used more for load balancing than for failover, problems like this will occur. (I used to work on failover drivers for a SAN).
- jackowayed 14y agoGitHub's failover problems have never been load-related. GitHub has pairs of fileservers where one is the master and the other's sole job is to follow along with the master and take over if it thinks the master is down, so when they do failover, it is to a node with just as much capacity as the previous master. All the failover problems I can think of since they moved to this architecture 4 years ago have been coordination problems where something undesired happened when transitioning from one member of a pair to another. In this case, network problems lead them to a state where both members of a pair thought they were the master.
- regularfry 14y agoThe same argument applies to UPSes. I hear about far more outages caused by UPS failures than by PSU or supply failures.
- rdl 14y agoI've never seen a redundant PSU be worse than a single PSU, actually (like the dual-line-cord modules on many servers or network devices). PSU failures have gone way down in the past 15 years or so that I've been observing them. The only power supplies I routinely see dying are external transformers on low end network devices and on systems exposed to really dirty power. I have seen facility-scale UPSes go bad, and sometimes in weird ways, but an order of magnitude less frequently than grid power. I think we reached a crossover point where designing for facility-scale survivability vs. replicating facilities ceased to be worthwhile for most Internet applications sometime in the past 10 years. It doesn't really make sense to drop $2b on a ~100k sf datacenter like AboveNet used to do for e.g. 365 Main. There are still some systems where replication is a pain, but even for those, I think metro area replication shouldn't be that hard. Even just running 5km of fiber in a loop between a few buildings in the same town gets you a huge amount of resilience against most facility problems.
- gizmo686 14y agoEvery other issue automaticly fails over.
- ghshephard 14y agoThis may sound selfish, but github does such a great job of writing up post mortems, that I almost look forward to their outages just because I know I'm going to learn a lot when they write their follow up.
- tptacek 14y agoCame here to say the same thing. This is Mark Imbriaco's wonder twin power. It really is something to generate goodwill from an outage writeup. Part of it is just unflinching transparency combined with nerdy details; you never feel like they're hiding anything, and you get to learn about all the operational doodads they're working with to run at this scale.
- imbriaco 14y agoThanks, I really appreciate that. For me, the motivation for transparency came from too many frustrating instances of being kept in the dark after things had gone wrong. The worst thing both during and after an outage is poor communication, so I do my best to explain as much as I can what is going on during an incident and what's happened after one is resolved. There's a very simple formula that I follow when writing a public port-mortem: 1. Apologize. You'd be surprised how many people don't do this , to their detriment. If you've harmed someone else because of downtime, the least you can do is apologize to them. 2. Demonstrate understanding of the events that took place. 3. Explain the remediation steps that you're going to take to help prevent further problems of the same type. Just following those three very simple rules results in an incredibly effective public explanation.
- daeken 14y agoThis sort of approach is the reason that when I need to upgrade to a higher plan on Github, I don't flinch. In fact, I love giving you guys more money, simply because you make my life completely painless; I don't think I can say the same about any other service. Keep up the awesome work.
- thaumaturgy 14y ago
- deleted 14y ago[deleted]
- apeace 14y agoIt's about time we heard from them. I understand the timing of this was unfortunate (right before a holiday), but the trust alluded to in the conclusion of that post would be bolstered by faster post-mortems on major outages like this one.
- deleted 14y ago[deleted]
- nixgeek 14y agoBetter to take a few days to gather evidence and really understand the problem, than rush out a post-mortem which is inaccurate or incorrect leaving the impression that downtime is not taken seriously and investigated thoroughly. IMO.
- InclinedPlane 14y agoI disagree. Other than for curiosity needs what's the value in a faster post mortem? Especially considering that a faster post-mortem would likely be a less accurate and less complete post-mortem? What is the difference in actionability for anyone between an update in 2 days vs an update in 4 days? I see none.
- apeace 14y agoTo me, a post-mortem doesn't just satisfy curiosity, it eases fears that the problem will return and informs me of future plans which may help prevent the problem, or may bring it back. It helps me form my own plan, since I'm a user of Github. I for one spent my break hawking my email, in case further Github outages caused any of my automated deploy scripts to fail. I realize it's my own responsibility to write scripts that handle failure scenarios, but the fact is my company pays Github to host our repositories. Downtime happens, I'm understanding of that. But when it does, I want to know what's going on as soon as that information's available--especially when I'm on holiday. I don't think a blog post written after they resolved all the issues would have been less accurate. It just would have inconvenienced whoever wrote it on a holiday. Not the biggest deal in the world--it would take me a lot more than this slip-up to switch from Github. But IMO a service provider should get at least some information out faster than this.
- deleted 14y ago[deleted]
- mememememememe 14y agoI am not trying to be hash. But again? How many more github downtime posts do I have to read on HN every month?
- raverbashing 14y agoAnd High Availability isn't. Again It seems redundancy protocols end up grappling each other more often than not Unfortunately, there is no easy answer, and I'm sure Github employs people with lots of experience. This makes me wonder about after several people working on problems like this it's still a challenge
- ChuckMcM 14y agoYou can't predict what you can't predict. This is what makes the Chaos Monkey experiment so interesting. And yes, HA is hard, and with many things it is hard with respect to latency. The more latency you can tolerate, the easier HA becomes. At an extreme, if you can tolerate 1 minute latency than each request can come in and compute a most likely way to complete a request with the most authoritative set of actors. Few people though are willing to tolerate a commit taking 15 minutes, much less a couple of hours. This was one of the most insightful things about NFS and the whole stateless design. By burdening the client with the state the server could be much simpler. Once you get above a certain size the problems change becoming both easier and harder (easier in that you can disperse your data further, harder in that your confidence in agreement between copies takes longer to compute and thus increases latency). It would be interesting if Google shared their work on Spanner (they seemed to have tackled this problem at a large scale) and given Netflix's experience (Chaos Monkey's dad) it seems like Amazon still hasn't quite gotten the recipe right. It is a deliciously thorny problem with subtle complexity and unexpected inter-dependencies.
- onetwothreefour 14y agoAhhh... good old Heartbeat. We used to use Heartbeat in a similar setup back in 2001. It was the worst architectural decision we ever made, and after one too many a failure (where STONITH/split-brain/etc killed the wrong machine, or both machines) we threw it out. TL;DR: This will happen again. Guaranteed.
- cookiecaper 14y agoWhat do you suggest as a replacement? Heartbeat is in use at many big players.
- mitchty 14y agoI'll just point out that experience with heartbeat from 2001 is rather outdated. That would be similar to comparing 2.4 kernels problems to the most recent 3.7 kernel. Likely not an overly useful anecdote. For the record we use heartbeat at work with no issues such as this.
- onetwothreefour 14y agoWhile I agree with you, I think the anecdote is still useful because it shows that the problems of the present are really problems of the past too. Software quality improves, sure, but you can still learn from the past.
- mitchty 14y agoI'd say its more an implementation issue than software. And that this is a solved issue even 10+ years ago with heartbeat. Our heartbeat links are segregated from the public network with separate network cards/switches. So the specific issue github hit here isn't what we would have encountered. We do have quad nic cards for a reason in the systems we run, this issue github hit is one of the various reasons you don't run heartbeat over the public topology. It will bite you in the ass no matter what cluster software you are running. Unless you also have a disk heartbeat over shared fibre/scsi, or maybe serial but same difference. Depending upon the public network though is a lost cause.
- ewokhead 14y agoNote to Github: Freeze prod changes two weeks before and two weeks after all major holidays. Your employees probably don't appreciate the hassle when all they are thinking about is "YEAH! DAYS OFF!" Just my opinion and how I run my systems in the DC.
- anu_gupta 14y agoI guess one of the counter-arguments to this (very good) suggestion, is that holidays are probably the quietest time in terms of traffic and usage.
- dlisboa 14y agoStopping coders from deploying stuff due to a risk of an unspecified "something" going bad makes for frustrated employees. By that time you might as well give them the days off as they'll have little to strive toward. You don't need to freeze all code deployments and other things that have little risk. Also, they most likely scheduled this at this time due to the lower traffic, probably the lowest of the year for them. While half of Github was probably enjoying their families the other half was planning for this for a long time. Given the size of the operation I don't think anyone took it lightly.
- ewokhead 14y agoSwitch modifications should never stop work. Prod code push != prod infrastructure changes. Which is what the article is talking about. Specifically the agg. switching layer. My reply is not about code deployments. It is about managing network devices with high visibility and impact. I still stand by my original comment with a critical detail added: Freeze prod ~infrastructure~ changes two weeks prior and two weeks after major holidays. Push code all you want. The RFO that they provided addresses link aggregation changes which are a part of an infrastructure change.
- nixgeek 14y agoHolidays are actually one of the best times to be making changes as traffic is significantly lower, and IMO, one should be aiming for an infrastructure where you can always ship changes without being afraid of the ramifications. Architecturally that may mean many things - hitting "SHIP IT!" might push code into a staging environment for some final testing before delivering it onto a platter in production. Should you have multiple sites, it might involve rolling out the new stuff to just one of them until you see how it goes. Maybe you have feature flags and want to introduce a new change to all servers, but just 1% of the user population? Fundamentally hitting "SHIP IT!" should be doing just that. Any constraints you put on how fast it gets to 100% of the user population are a risk control, and you need to optimize for a balance of developer happiness and system stability. When you concede "We can't make changes because we're frozen" outside of a critical systems ('life critical') environment, you should quit your IT job and go become a fisherman or something.
- cbsmith 14y agotl;dr: It's really hard to get high availability systems right, and we still run the entire service as a single colo. I can totally understand this kind of thing going wrong, but particularly given the service they provide, why not have a second colo, with a relatively recent clone of the repo, that you can route people to? Heck, you can likely even do an automatic merge once the other repo is working again...
- MichaelGG 14y agoRelatively recent clone? Sounds like that would screw customers up pretty bad if they don't realise the problem. If they went to a second site, having synchronous commit to both sites is how it should be done, no? The extra latency on infrequent git pushes is far less an inconvenience than the possibility of grabbing the wrong code.
- cbsmith 14y ago> Sounds like that would screw customers up pretty bad if they don't realise the problem. I think there'd be a variety of ways to have the system to fail until the customer made some kind of adjustment that indicated they grokked that there was a failure (like say... changing your upstream).
- dos1 14y agoWho's their network switch vendor? I'm not a networking expert, but boy - it sure seems like their switch vendor has screwed some things up royally. Or perhaps this is common with all complex network topologies regardless of hardware vendor? EDIT: I would just like to say, along with others, I greatly enjoy their postmortems and I feel as though I learn something every time. Kudos to them for being forthright. I host my personal and professional projects with them and am supremely confident that my data is as safe with them as it is with anyone.
- rdl 14y agoThey are hosted at Rackspace, so I think the default is "Vendor C". I'm curious if they'd even be able to support an Arista network. Although this would probably not be a problem with that kind of network.
- sounds 14y agoThey don't use Cisco for their Aggregation switches.
- rdl 14y agoWhen did that change? They did when I last checked (which, admittedly, was Some Years Ago.)
- nixgeek 14y agoRecently.
- rdl 14y agoTo what, some other vendor for price? (oh, I see you are with GitHub. Somehow I suspect you did not push for this downgrade.) (followup to sounds: I'm not particularly pro cisco except that all-cisco lets you avoid interop problems. Since at least some of this is cisco, there's a good argument for all cisco. I've actually used HP and Dell successfully for certain things, but only because I kept it as simple as possible. In the past using other stuff was necessary at the high end, too, because cisco didn't provide 10GE very well, etc. What I am against is cheaping out on your ~few aggregation switches, which I've seen other people do, when you're already spending good money on everything else. Maybe it's different once you get to Google or Facebook scale, but I've never been responsible for quite that many switches, and an extra $500-$1000 per 24 ports isn't going to kill you. I actually prefer non-cisco substantially for routing and for security products.)
- mikec3k 14y agoI always love to read post-mortems like this. It's fascinating to see how a simple event can trigger a massive failure & we can learn a lot from them.
- chuhnk 14y agoGitHub as a sysadmin/system engineer I feel your pain, I understand it completely and know the horrors of failures leading into multi-hour recovery. That said, you need to do better. This is a heavily relied upon resource for the open source community. Perhaps you didn't estimate this sort of growth but now you are here I'm sorry but the weight does fall on your shoulders. I heavily commend you guys on the service you've provided thus far and what you've done to actually pull all the varying language/library communities together. I honestly want to see this scale and serve 5 9s year round. You need to take a good hard look at the architecture of the stack and find a way to get it multi-homed.
- thaumaturgy 14y agoPlease don't take this personally, but I'd like to propose a new rule of etiquette on HN: it's unacceptably lazy to say that some business "needs to do better" without explaining exactly what they should do better, or how. GitHub is doing better by proactively upgrading their network. During the upgrade, they have run in to some technical difficulties. There are maybe, what, a thousand people in the U.S. that have worked on network administration for something GitHub's size? And far fewer who could have predicted this kind of trouble? If we're going to shit on someone, we should at least make it nutrient-rich: include recommendations based on experience from dealing with problems of that nature and magnitude.
- dsl 14y agoI understand they have a lot of nerd/hacker cred in and around the tech hotspots of the country, but GitHub is by no measure "big". The most recent number I could find puts them at 33 servers total. I've worked for non-tech companies that have that many as hot standbys. One big recommendation I'd offer is to have a secondary network with as few moving pieces as possible (think a dumb unmanaged Netgear switch or two) to run things like heartbeats and DRBD over. That sort of thing should not be living on the same segment as your production traffic.
- imbriaco 14y agoWe have significantly more than 33 servers.
- robomartin 14y agohttp://www.youtube.com/watch?v=c8N72t7aScY http://www.youtube.com/watch?v=c8N72t7aScY
- el_cuadrado 14y agoHigh Availability strikes again. No surprise there. But I am mildly flabbergasted by the fact that GitHub uses STONITH. The technology is as safe as open-core nuclear reactor, and it works reliably only in very simple conditions.
- deleted 14y ago[deleted]
- mofraw 14y agoSTONITH is designed for critical failure so you don't end up with a split-brain situation, which is far worse than a dead node. STONITH is a good thing. The problem here is more than cluster wasn't configure to survive a catastrophic switching failure.
- el_cuadrado 14y agoYep, right - fencing solution that relies on the network to stop the service. How could that possibly fail?
- ewokhead 14y agoI just realized that the Sys. Admin/Prod. Ops to Developer ratio here is crazy low. Everyone assumes I am talking about code changes when the article is about prod switching and network transit device changes. MLAG or any LAG technology, LACP, bonds whatever should never impact the deployment of code. It should be invisible when it is working. Obviously it is very visible when it breaks though. My heart goes out to the Github guys! Sorry for the confusion everyone.
- jetsnoc 14y agoWow, I'm very glad our company chose a routed design with an interior routing protocol (OSPF.) I've never been able to push the limits of a layer two network as far as GitHub. A routed network helps segment the network so when systems fail or a re-convergence mistakenly occurs only a few racks are having problems and not the entire system. It's also very helpful for us to push routes to our exterior routing protocol (BGP.) I also find it interesting they don't use any additional out-of-band network for heartbeats/management especially as unstable as their layer two network has been. It sounds like file servers need a secondary stable heartbeat network even if it is only 10/100. No judgements being passed here it just seems like a lot of eggs in one basket. That being said, Thank you for this write-up and sharing so openly and honestly. Happy GitHub customer here! EDIT: Yes, I know routed networks can have similar problems but they are designed for routing, pathing and redundancy with a lot less overhead on the broadcast domain.
- imbriaco 14y agoYou're right on all counts. We have a great many plans with regard to how we want our network to operate that are underway.
- mprovost 14y agoRouted networks tend to fail closed where switched networks tend to fail open. That usually seems to be the problem with layer two failures, they can easily spread traffic everywhere and it takes a while for things to calm down, in the meantime the network hardware is often overwhelmed. With layer three networks the failures tend to be that you lose the ability to route somewhere but in my experience it's often easier to recover from that situation.
- akg_67 14y agoI get the impression that issue wasn't network hardware but bad high availability design on fileserver side. Why do GitHub has failover network on same network hardware as primary network? But I am not surprised as I see this at lot of clients that they have failover network on separate VLAN on same network hardware. And, whenever they have network hardware issue, servers run into split brain problems. The failover network should be totally different physically and logically from primary network. The heartbeat between file servers should be checked through both primary network and failover network. If a server can't be reached by its partner over primary network, it should be gracefully taken offline by partner through failover network.
- ksec 14y agoI have absolutely ZERO knowledge on Enterprise Networking. But it strikes me that something as dump as Router, Switches or Network Equipment are still so unstable. A Hypothetical questions, why not something a 8 ARM 64 bit core Linux computer as switches? Making more logic resides in software instead?
- regularfry 14y agoThere is a move to put more of the switching and routing logic in software (google for openvswitch if you're interested) but part of the problem is that general purpose hardware doesn't stand a hope in hell of keeping up with interesting network data rates. You absolutely need to be doing a fair portion of the work in hardware, the CPUs just coordinate it.
- nixgeek 14y agoIt's probably worth reviewing this article on Intel FM6000 SDN to get a glimpse into the complexity involved. http://www.intel.com/content/dam/www/public/us/en/documents/white-papers/ethernet-switch-fm6000-sdn-paper.pdf http://www.intel.com/content/dam/www/public/us/en/documents/...
- treskot 14y agoI've noticed MC-LAG / MLAG failing quite often in my case. Any details over the MLAG fail? Are we doing it wrong? Alternatives?