24 ms·
January 28th Incident Report
- cognivore 11y agoUm, work from your local cache for a few hours? It's that the one of the main reasons for git?
- majewsky 11y agoNot all processes that involve GitHub are development processes. I've seen automated deployments fail inside a corporate network when the resident HTTP proxy had a bad day and could not connect to github.com.
- bosdev 11y agoThere's no mention of why they don't have redundant systems in more than one datacenter. As they say, it is unavoidable to have power or connectivity disruptions in a datacenter. This is why reliable configurations have redundancy in another datacenter elsewhere in the world.
- emergentcypher 11y agoSeriously. I'm kind of surprised about this.
- cookiecaper 11y agoYeah, they gloss over it but at its heart, keeping mission-critical servers in a single datacenter with no redundancy is among the most common and amateur infrastructure failures. Many would expect a company like GitHub to have anticipated and prevented it. GitHub should have a process to ensure that all services are redundant before they get pushed to production.
- nemothekid 11y agoGiven the dependency in question is Redis, such a solution is probably exacerbated by the fact Redis hasn't really had a decent HA solution. This is also hidden by the fact that Redis is really reliable (in my experience at least). In my experience it usually takes an ops event (like adding more RAM to the redis machine) to realize where a crutch has been developed on Redis in critical paths.
- yeukhon 11y agoA lot of tools and services people use either don't have HA at all or don't have a native support for true distributed HA. But that can't stop people from making some HA or alike solution. I am not sure what they use Redis for but along the line of caching and key-value store they must have figured out how to invalidate data, otherwise they'd be running only a single instance of Redis. i.e. they are running "HA" just in a single data center, so logically speaking that's not difficult to port over to another data center.
- mwpmaybe 11y agoI'm not familiar enough with Redis's clustering features to speak to the exact issues with what you're proposing, but generally speaking, HA is almost a completely different problem than disaster recovery (DR). Sure, the protocol is the protocol, but you wouldn't want to cluster local and remote nodes together for several reasons, primarily latency, security, and resiliency. Performance will suffer if they're clustered together and a single issue could take down nodes in both data centers, which kind of defeats the purpose. What you really want is a completely separate cluster running in a different data center (site). It should be isolated on its own network and ideally it should have different admin rights/credentials and a different software maintenance (patching) schedule. A completely empty site isn't much use so you'll need some kind of replication scheme. Naturally, these isolating steps make site replication difficult. You might patch one site and now the replication stream is incompatible with the other site. (You can't patch both sites at the same time because the patch might take down the cluster.) Or whatever you're using to replicate the sites, which has credentials to both sites, breaks and blows everything up. You need a way to demote and promote sites and a constraint on only one site being the "master" at a time. What happens if network connectivity is lost between sites? What happens if one site is down for an extended period of time? Maybe you need a third, tie-breaking site? Once you work through these issues, you are still exposed to user error. Your replication scheme might be perfect... perfect enough that that an inadvertently dropped table (or whatever) is instantly replicated to the other site and is now unrecoverable without going to tape. Maybe you introduce a delay in the replication to catch these oopsies, but now your RPO is affected. Anyway, it's a bit of a shell game of compromises and margins of error. Source: 10 years designing and building HA/DR solutions for Discover Card.
- theptip 11y agoIt's shocking that they don't at least have a read replica of their system in another 'AZ'. That's cloud hosting 101, and being self-hosted isn't an excuse to skimp on this. If an outage caused 2 hours of read-only access to repos it would still be moderately impactful, but at least we could still build our Go code.
- drdrey 11y agoFor people reading this, AZ in this context would be Availability Zone
- paulddraper 11y agoRight, and not the Grand Canyon State. The space of acronyms/abbreviations is quite cluttered.
- theptip 11y agoHeh, not from the US initially so that overloading did not occur to me :) That's an interesting alternate reading though... "To ensure the integrity of our data, we need to locate another Arizona, since this one is serving us so well."
- deleted 11y ago[deleted]
- paulddraper 11y agoRight, and not the Grand Canyon State. The space of acronyms/abbreviations is quite cluttered.
- vacri 11y agoFor people reading this, 'Availability Zone' in this context would be (AWS speak for) 'datacentre'. :)
- NetStrikeForce 11y agoSo your building process depends on the availability of an external company?
- beachstartup 11y ago> There's no mention of why they don't have redundant systems in more than one datacenter sometimes reading comments on hn makes me laugh out loud. there's only one reason to not do this, and that's cost. what do you expect them to say about that? i mean really, you think they're going to put that in a blog post: "Well, the reason we don't have an entire replica of our entire installation is because it costs way too much. In fact, more than double! And so far our uptime is actually 99.99% so there's no way it's worth it! You can forget about that spend! Sorry bros."
- menssen 11y agoThis is not only obviously true, I think it is also a completely reasonable calculus. They just proved that if the entire Redis cluster goes down they can get it back in 2.5 hours. It's almost certainly a caching layer, so there is no permanent data loss. If they fix the application bootstrap dependency on a Redis connection, and they add monitoring to more easily see in the future when the Redis cluster is the problem, next time that time period will probably be way shorter. So, a very small risk of an hour or so of downtime sometime in the future which will not cause data loss, or tens of thousands of dollars a month for a failover cluster? I wouldn't replicate it either.
- cookiecaper 11y ago>It's almost certainly a caching layer, so there is no permanent data loss. People who use Redis rarely end up using it solely as a caching layer. It often also takes on the role of an RPC facilitator and pseudo-database. GitHub's post also mentions that their engineering team had to replicate Redis' dataset before they could get the alternative hardware running, which implies that they do need some data in there before the site is operational. Personally one of my pet peeves is people throwing mission-critical data in Redis and acting like it's honky-dory. It happens all the time and seems really difficult to get people to not do. There's a reason we have a real ACID compliant database storing non-disposable data; it's ridiculous to ignore that just because it's easier to stuff it in Redis. I think it's reasonable to have a dependency on a Redis server, but I don't think it's reasonable to depend on any data in particular being stored in that server. It should be used as a caching/acceleration layer for data that can be easily and automatically regenerated.
- mjevans 11y agoThis just shows how difficult it is to avoid hidden dependencies without a complete, cleanly isolated, testing environment of sufficient scale to replicate production operations and do strange system fault scenarios somewhere that won't kill production.
- imbriaco 11y agoIt turns out that it's even hard then. Complex systems, by their very nature, fail in unexpected and unpredictable ways. If that weren't bad enough, hindsight bias makes it way too easy for us to look back with perfect knowledge and opine "That was so obvious, how could they have missed such a rudimentary issue?" If only things were that easy.
- ssmoot 11y agoI'm not sure what part of servers failing to POST is especially complex or related to distributed computing. For all the fawning over being provided technical details, this article was pretty light on them. I don't think Github going down for a couple hours is that big of a deal TBH. But it does seem to expose a few really basic failings in their DR planning IMO. I also think it's ridiculous that some commenters are trying to frame this as a distributed computing problem. It's not even a clustering problem (apparently). It's just looking at the iDRAC or whatever to see why the server isn't getting past POST and putting your recovery plan into action. This is white box vanilla stuff that happens to everybody. That servers had to be rebuilt as part of DR says a lot. The fact that there was a Redis dependency during bootstrap? Probably a good thing. You know as well as anyone I'm sure the last thing you want is a bunch of processes that only look like they're up. And even if they could not error without their Redis connections, if Redis is used for caching, what's that going to do to availability? Would it be a good thing to have the processes up if they can only handle 10% of the usual load? Those are details that aren't there. But complex distributed computing problem this is not. Not as it was presented anyways.
- ones_and_zeros 11y agoOr use the Netflix model: Chaos testing in production.
- merqurio 11y agoI feel it was good incident for the Open Source community, to see how dependent we are on GitHub today. I feel sad whenever I see another large project like Python moving to GitHub, a closed-sourced company. I know, GitLab is there as an alternative, but I would love to see all the big Open Source projects putting pressure over GitHub to make them open their source code, as right they are big player in open source, like it or not.
- davidcelis 11y agoWas it a good incident to see how dependent we are on GitHub? Every time there's a GitHub outage, a vocal group of people will voice their opinions that we are too dependent on GitHub, we should be using open source alternatives, GitHub should be open source, etc. Then, within a few days, everybody goes silent and we return to our normal lives. I don't think outages at GitHub are very frequent. This one was lengthy, so it's definitely been on a lot of peoples' minds, but this conversation always comes up when it happens.
- merqurio 11y agoOf course it was. I don't know if everybody goes silent after a few days, It's the first outage I'm aware of, but some people at university made see the hypocrisy of using GitHub for open source projects and I feel that if there is a community strong enough to make some impact on GitHub that could be hackernews. Maybe I'm wrong.
- davidcelis 11y agoIf you look back through the years and find a few other stories of "GitHub is down", you'll see that this conversation happens every time. Some people tread into the HackerNews thread and say "More people should be using self-hosted GitLab instances" or "if GitHub would just open source their code, we wouldn't need to be so dependent." But then the conversation stops within days because, the fact is, hosting your own git servers and getting people to actually use them is a huge pain in the ass. More simply put: people just like using GitHub. Furthermore, GitHub's a business. They're selling private repositories. They do open source quite a bit of code, but they're not going to open source their actual product.
- dsmithatx 11y agoIf only Bitbucket could give such comprehensive reports. A few months back outages seemed almost daily. Things are more stable now. I hope for the long term.
- viraptor 11y agoIsn't BB's problem basically that there are too many users? GH's outage writeup is cool, because it's a one off and it can be analysed. When BB is just overloaded for a long time and needs more power, it's not going to be very interesting. (unless I missed some specific non capacity related outages?)
- yeukhon 11y agoMaybe. BitBucket was also an acquisition so for some time I believe there was a lack of resource provided to them and there was a huge technical debt/integration effort required. At this very time, I don't know if Atlassian actually care much about BitBucket. They are probably more concerned about delivering Stash than BitBucket, my wild guess. I was an active BB user a couple years ago, and the project I worked on would hg clone from BB many times a day so I would be the first one to notice a 503 or whatever error coming from their service. Typically I would see one or two outage per month, some last a few minutes, some last several hours. Most of the time the outage impacted git/hg checkout, so I think that was their technical bottleneck.
- lhc- 11y agoFYI, Stash is now Bitbucket Server, and the plan as I've heard it is to work towards feature parity between the two.
- mattdeboard 11y agoWe use Stash and it is surprisingly not bad at all. Github is much more polished but for code browsing and review, it does that which it is supposed to do.
- 11y ago
- danielvf 11y agoFor all that work to be done in just two hours is amazing, especially with degraded internal tools, and both hardware and ops teams working simultaneously.
- imbriaco 11y agoYou're absolutely right. We should collectively be using incidents like this as an opportunity to learn, much like the GitHub team does. Our entire industry is held back by the lack of knowledge sharing when it comes to problem response and the fact that so many companies are terrified of being transparent in the face of failure. This is very well written retrospective that gives us a glimpse into the internal review that they conducted. Imagine how much we could collectively learn if everyone was fearless about sharing.
- totally 11y agorelevant: https://codeascraft.com/2012/05/22/blameless-postmortems/ https://codeascraft.com/2012/05/22/blameless-postmortems/
- ssmoot 11y agoIs there a timeline to how long it took them to figure out Redis was down? Because having experienced the same, you get an alert. Cool. HA-Proxy says app servers are down. Ok. You SSH in and see that everything looks ok but the processes are bouncing. You tail the logs to find out why (obviously lots of these steps could be optimized). Within a few seconds you spot the error connecting to Redis. A minute later you've verified the Redis hosts are offline. That's the first 5 minutes after getting to a computer. After that it doesn't really matter why they're down. You failover, get the site back up and worry about it later. Are these systems on a SAN? That's probably the first mistake if so. Redis isn't HA. You're not going to bounce it's block devices over to another server in the event of a failure. That's just a complex, very expensive strategy that introduces a lot of novel ways to shoot yourself in the face. If you're hosting at your own data-center, you use DAS with Redis. Cheaper, simpler. I've never seen an issue where a cabinet power loss caused a JBOD failure (I'm sure it happens, but it's a far from common scenario IME), but then again, locality matters. Don't get overly clever and spread logical systems across cabinets just because you can. Being involved with this sort of thing more frequently than I'd like to admit, I don't know the exact situation here, but 2h6m isn't necessarily anything to brag about without a lot more context. What's pretty shameful is that a company with GitHub's resources isn't drilling failover procedures, is ignoring physical segmentation as an availability target (or maybe just got really really unlucky; stuff happens), and doesn't have a backup data-center with BGP or DNS failover. This is all stuff that (in theory if not always in practice), many of their clients wearing a "PCI Compliant" badge are already doing on their own systems.
- pedalpete 11y agoDoes Github run anything like Netflix Simbian Army against it's services? As a company by engineers for engineers with the scale that github has reached, I'm a bit surprised they are lacking a bit more redundancy. Though they may not need the uptime of netflix, an outage of more than a few minutes on github could affect businesses that rely on the service.
- deleted 11y ago[deleted]
- deleted 11y ago[deleted]
- imbriaco 11y agoGoogle "Netflix downtime" for evidence that Netflix also has outages. Google has outages, sometimes very significant ones of Google Apps. Facebook has outages. Complex systems fail. Period. All the time. Things like the Simian Army are fantastic tools that help you identify a host of problems and remediate them in advance, but they cannot test every combinatorial possibility in a complex distributed system. At the end of the day, the best defense is to have skilled people who are practiced at responding to problems. GitHub has those in spades, which is why they could respond to a widespread failure of their physical layer in just over 2 hours. The biggest win with the Simian Army isn't that it improves your redundancy. It's that it gives your people opportunities to _practice_ responses.
- drdrey 11y agoMore than practicing responses, Chaos Monkey and Failure Injection Testing allow us to verify that we don't have unexpected hard dependencies. Sometimes you find out that your service can't start if another one becomes latent, in which case you can plan for it by adding redundancy/extra capacity, fallbacks or working in degraded mode.
- kuschku 11y agoI remember in 2013 a full-day outage of Google.
- onetwotree 11y agoEvery time I read about a massive systems failure, I think of Jurassic Park and am mildly grateful that the velociraptor padock wasn't depending on the systems operation.
- chris_wot 11y agoI think you'll find they were.
- mattdeboard 11y agoWell as long as you're not Samuel L. Jackson in that scenario you should be fine. Ish.
- onetwotree 11y agoSamuel L. Jackson taught me everything I know about ethics in software engineering. Including the principle that if your software breaks, you're the on who has to go get savaged by velociraptors to fix it.
- swrobel 11y agoAnyone got a good tl;dr version?
- alblue 11y agoPower outage in DC brought many machines down. Redis clusters failed to start owing to disk issues (not cleanly unmounted?). The reboot of remaining machines uncovered an unknown dependency on the machines needing the redis cluster to be up in order to boot. There were other learning points such as immediately going into anti DDoS mode and human communication issues that didn't realise or escalate the problem until some time after the issues started occurring.
- daigoba66 11y agoLost power. Took a while to get the servers cleanly rebooted.
- maerF0x0 11y agoIntern trips on power cable, 25% of servers go down. Edit this is mostly the "DR" part of tldr :P
- aidenn0 11y agoPower outage brought 25% of servers down. Firmware issue meant that a large fraction of their servers could not detect the disks on reboot. This prevented the redis cluster from starting. They inadvertently have a hard-dependency on redis being up for the majority of their infrastructure to start.
- draw_down 11y ago"Stuff went wrong and our servers were down for a couple hours." You're welcome.
- contingencies 11y agoNo CI/test process was in place for critical systems to ensure that they had no external dependencies. Takeaway: If you run any complex system, ensure that each component is tested for its response to various degrees of failure in peer services, including but not limited to totally unavailable, intermittent connectivity, reduced bandwidth, lossy links, power-cycling peers. No CI/test process was in place for hardware/firmware combos to ensure they recovered fine from power loss. Takeaway: If you run a decent-sized cluster, ensure all new hardware ingested is tested through various power state transitions multiple times, and again after firmware updates. With software defined networking now the norm, we have little excuse not to put a machine through its paces on an automated basis before accepting it to run critical infrastructure. No CI/test process was in place for status advisory processes to ensure they were sufficiently rapid, representative, and automated. Takeaway: Test your status update processes as you would test any other component service. If humans are involved, drill them regularly. Infrastructure was too dependent on a single data center. Takeaway: Analyze worst case failure modes, which are usually entire-site and power, networking or security related. Where possible, never depend on a single site. (At a more abstract level of business, this extends to legal jurisdictions). Don't believe the promises of third party service providers (SLAs). PS. I am available for consulting, and not expensive.
- viraptor 11y ago> ... Updating our tooling to automatically open issues for the team when new firmware updates are available will force us to review the changelogs against our environment. That's an awesome idea. I wish all companies published the firmware releases in simple rss feeds, so everyone could easily integrate them with their trackers. (If someone's bored, that may be a nice service actually ;) )
- Cthulhu_ 11y agoI've played with the idea of some automated software update reporting site ages ago - it'd read rss feeds and scrape websites for the required info. It'd probably need adjustments for each hardware manufacturer / product though, and regular updating. But that could possibly be part of an open source project, give the firmware maintainers the opportunity to help out too.
- vhost- 11y agoThis was one of the toughest things about admining hardware clusters. Firmware updates (and firmware issues) are so hard to track down. It's so annoying. I remember spending a week tracking down an issue with a RAID controller and then spending another day or two on the phone with the vendor trying to get a firmware update so we did not have 2 racks of hardware sitting on a ticking time-bomb.
- DarkTree 11y agoI don't know enough about server infrastructure to comment on whether or not Github was adequately prepared or reacted appropriately to fix the problem. But wow it is refreshing to hear a company take full responsibility and own up to a mistake/failure and apologize for it. Like people, all companies will make mistakes and have momentary problems. It's normal. So own up to it and learn how to avoid the mistake in the future.
- eric_h 11y agoAs I said in another comment, the fact that they found an 8 minute delay from outage to status page update to be unacceptable speaks volumes to how much they value their relationship with their customers. as an aside I feel that I'm quite fortunate to work in the EST timezone, as their outage apparently started at about 7pm my time. We have a general rule at my company to not deploy after 6pm unless an emergency fix absolutely needs to go up. I saw the title of the story and said to myself, what outage? :P
- totally 11y ago> Because we have experience mitigating DDoS attacks, our response procedure is now habit and we are pleased we could act quickly and confidently without distracting other efforts to resolve the incident. The thing that fixed the last problem doesn't always fix the current problem.
- dgritsko 11y agoOccam's razor isn't a bad rule of thumb, however.
- jargonless 11y agoWhat is this "HA" jargon? I would STFW, but searching for "HA" isn't helpful.
- polysaturate 11y agoPretty sure it's "High Availability" in this instance...
- suraj 11y agoHigh availability
- unsatchmo 11y agohttp://lmgtfy.com/?q=HA+Server http://lmgtfy.com/?q=HA+Server
- deleted 11y ago[deleted]
- mattbeckman 11y agoAlso, of note, the ever popular HAProxy Github uses and mentions stands for "High Availability Proxy".
- xzlzx 11y agoYou could google "HA", click in the Wikipedia link that shows all the things "HA" may refer to, and deduct that the most logical thing in the list, given the context, would be this link: https://en.wikipedia.org/wiki/High_availability https://en.wikipedia.org/wiki/High_availability.
- tonylxc 11y agoTL;DR: "We don’t believe it is possible to fully prevent the events that resulted in a large part of our infrastructure losing power, ..." This doesn't sound very good.
- jrockway 11y agoIf your plan to avoid downtime is to prevent power outages, you're going to have downtime. All their sentence says is they can't prevent power outages. That's fine, because the other 1/nth of your servers are on a different power grid in a different state.
- tonylxc 11y agoI totally share the same view that to best avoid failure is to embrace it and cope with it. It is true that all their sentence is about recovery, however, it is disappointing that they didn't mention anything about a redundant datacenter.
- deleted 11y ago[deleted]
- theptip 11y agoThe rest of the sentence is pertinent: "...but we can take steps to ensure recovery occurs in a fast and reliable manner. We can also take steps to mitigate the negative impact of these events on our users." The lessons that giants like Netflix have learned about running massive distributed applications show that you cannot avoid failure, and instead must plan for it. Now, having a single datacenter is not a good plan if you want to give any sort of uptime guarantee, but that's a different point to make.
- tonylxc 11y agoMy point is: they shouldn't ONLY plan on ensuring recovery occurs fast; they should also plan on having multiple data centers, which to me is more important. It's frightening to know that such an important service is only operating in a single data center. However, their recovery report didn't mention anything about such a plan. << Edited: correct a grammar error.
- mattdeboard 11y agoAnyone have a link to a description of the firmware bug that caused the disk-mounting failure after power was restored?
- ymse 11y agoI'm going to guess that these are Dell R730xd boxes with PERC H730 Mini controllers (LSI MegaRAID SAS-3 3108). A failed/failing drive present during cold boot could cause the controller to believe there were no drives present. To add insult to injury, on early BIOS versions this made the UEFI interface inaccessible. The only way to recover from this state was to re-seat the RAID controller. There were also two bizarre cases where the operating system SSD RAID1 would be wiped and replaced with a NTFS partition after upgrading the controller firmware (and more) on an affected system (hanging/flapping drives). Attempts to enter UEFI caused a fatal crash, but reinstall (over PXE) worked fine. BIOS upgrade from within fresh install restored it. From the changelog: Fixes: - Decreased latency impact for passthrough commands on SATA disks - Improved error handling for iDRAC / CEM storage functions - Usability improvements for CTRL-R and HII utilities - Resolved several cases where foreign drives could not be imported - Resolved several issues where the presence of failed drives could lead to controller hangs - Resolved issues with managing controllers in HBA mode from iDRAC / CEM - Resolved issues with displayed Virtual Disk and Non-RAID Drive counts in BIOS boot mode - Corrected issue with tape media on H330 where tape was not being treated as sequential device - resolved an issue where Inserted hard drives might not get detected properly.
- spydum 11y agoSo, while it sounds like they have reasonable HA, they fell down on DR. unrelated, I could not comprehend what this means?..: technicians to bring these servers back online by draining the flea power to bring Flea power?
- Someone1234 11y agoI assume they mean completely disconnect the equipment from ALL external power sources. Typically even when a piece of equipment is offline in a data center, it continues to draw power, and will often keep running systems like DRAC and other management/status tools (since the whole concept of a data center is NEVER having to get up out of your chair, so even a "shutdown" system needs to be able to be remotely started). Since the firmware had a bug, bad state could be stored, completely removing power may clear that state and appears to have done so in this case. They may have also needed to pull the backup battery, and reset the firmware settings, but I wouldn't presume that just from the term "flea power."
- spydum 11y agosure enough, it's a real term, and it's relatively old.. http://answers.google.com/answers/threadview/id/185999.html http://answers.google.com/answers/threadview/id/185999.html I have never known what to call this, but have definitely been engaged in draining a few fleas. Also, I can't believe it's been that long since google answers has been closed..
- tmsh 11y ago> Over the past week, we have devoted significant time and effort towards understanding the nature of the cascading failure which led to GitHub being unavailable for over two hours. I don't mean to be blasphemous, but from a high level, is the performance issues with Ruby (and Rails) that necessitate close binding with Redis (i.e., lots of caching) part of the issue? It sounds like the fundamental issue is not Ruby, nor Redis, but the close coupling between them. That's sort of interesting.
- lukeasrodgers 11y agoAs someone with a fair bit of ruby+rails+redis experience, I don't think this is blasphemous, but I also don't think the performance issues of ruby/rails having anything to do with the failure. Generally you would cache/store something in redis not because your programming language or framework is slow, but because a query to another database is slow (or at least, slower than redis), or because redis data structures happen to be a good/quick way to store certain kinds of data. I believe the fundamental issue was just that redis availability was taken for granted by app servers so that certain code paths/requests would fail if it wasn't available, rather than merely be slower.
- atom_enger 11y agoI don't think that Ruby/Rails has anything to do with this, really. If you want to scale any app, you're going to want to do some caching somewhere. What this boils down to is that their app has a dependency in an initializer that depends on redis. Without a connection to redis, it will flap.
- byroot 11y agoNo the fundamental issue is that an application should not require any external service to boot. It has nothing to do with Ruby, or Rails or even Redis. It's just a design flaw of the application, that you often learn the hard way.
- guelo 11y agoWeird that they didn't say what caused the power outage and what the mitigations are for that.
- sh4na 11y agoIf it's a data center owned by a third party, they probably can't talk about it.
- gsibble 11y agoI'm also confused about how the racks would lose power. Surely they had UPSes.
- ams6110 11y agoUPSs don't always cover everything. There are systems that are considered critical that are on UPS, and others that are considered restartable that might not be. There are a lot of tradeoffs in a data center. Having full UPS and generator backup capacity for everything gets very expensive.
- technion 11y agoI have multiple experiences with high end DCs with dual UPS and diesel genset experiencing power fail. Once it involved fire alarms, which trigger safety shutdowns within a suite. The other involved a failed static switch panel - ie, the things that aren't mean to be able to fail.
- abrookewood 11y agoGenerally speaking, I'd recommend AGAINST running UPSes in racks that are managed by top-tier data centres. I've had way more trouble with UPSes misbehaving than I ever have with data centres losing power. EDIT: I'd also point out that 2 hours is a long time to be running on in-rack UPSes. I've usually seen them designed to withstand about an hour, but not much more.
- WatchDog 11y agoThe power outage was only brief, enough to halt the servers but, much less than the 2 hour outage window.
- eric_h 11y ago> One of the biggest customer-facing effects of this delay was that status.github.com wasn't set to status red until 00:32am UTC, eight minutes after the site became inaccessible. We consider this to be an unacceptably long delay, and will ensure faster communication to our users in the future. Amazon could learn a thing or two from Github in terms of understanding customer expectations.
- dmunoz 11y agoI recently stepped into a role with a devops component, and one of my first surprises was just how slow status.aws.amazon.com was to update about ongoing issues. I had to scramble to find twitter and external forums confirmation for the client.
- atom_enger 11y agoWhat's even worse is that when Amazon finally updates their status page it's usually still a green icon with a little i tick for "information" even if it was a partial outage. It takes a lot for the icons to go red which is what you'd look for if you're experiencing issues. I do the same thing, often searching Twitter for "aws" or "outage" and find people complaining about the problem which confirms my suspicions. It's a sad state of affairs when you have to do this and Amazon doesn't seem interested in fixing it.
- click170 11y agoIf you have a support agreement with them then file a ticket requesting better customer communication and link back here as an example of how to do it right. I think everyone complains in forums and online but doesn't actually file tickets about it. These things are worth tickets too.
- eric_h 11y agoI think a lot of folks feel that it's a useless endeavor, so they don't bother. Amazon's been operating this way for years, and they're quite a large company; it seems unlikely to me that fundamental change can happen inspired by customer tickets, even if you're paying for support. Basically, if Netflix isn't the source of the complaint, they're not going to give two fucks. /me suspects that netflix engineers get outage notifications through some other avenue than the status page.
- TazeTSchnitzel 11y ago> We had inadvertently added a hard dependency on our Redis cluster being available within the boot path of our application code. I seem to recall a recent post on here about how you shouldn't have such hard dependencies. It's good advice. Incidentally, this type of dependency is unlikely to happen if you have a shared-nothing model (like PHP has, for instance), because in such a system each request is isolated and tries to connect on its own.
- rqebmm 11y agoIt must be nice to know that the majority of your customers are familiar enough with the nature of your work that they'll actually understand a relatively complex issue like this. Almost by definition, we've all been there.
- matt_wulfeck 11y ago> Remote access console screenshots from the failed hardware showed boot failures because the physical drives were no longer recognized. I'm getting flashbacks. All of the servers in the DC reboot and NONE of them come online. No network or anything. Even remotely rebooting them again we had nothing. Finally getting a screen (which is a pain in itself) we saw they were all stuck on a grub screen. Grub detected an error and decided not to boot automatically. Needless to say we patched grubbed and removed this "feature" promptly!
- rurounijones 11y agoI would have expected there to be a notification system owned by the DC that literally send an email to clients saying "Power blipped / failed". That would have given them immediate co text and not wasting time on DDOS protection
- deleted 11y ago[deleted]
- julesbond007 11y agoI seriously doubt this version of the story. While it's possible for several hardware/firmware to fail in all your datacenters, for them to fail at the same time is highly unlikely. This may just be a PR spin to think they're not vulnerable to security attacks. While this was happening at Github, I noticed several other companies facing that same issue at the same time. Atlassian was down for the most part. It could have been an issue with the service github uses, but they won't admit that. Notice they never said what the firmware issue was instead blaming it on 'hardware'. I think they should be transparent with people about such vulnerability, but I suspect they would never say so because then they would lose revenue. Here on my blog I talked about this issue: http://julesjaypaulynice.com/simple-server-malicious-attacks/ http://julesjaypaulynice.com/simple-server-malicious-attacks... I think it was some ddos campaign going on over the web.
- dandandan 11y agoThey're not hosted in multiple datacenters; there was a power interruption in their single datacenter that exposed this firmware bug. The point of this postmortem isn't the initial power interruption but rather its repercussions, why it took so long to recover from and how they can improve their response and communications in the future.
- julesbond007 11y agoOk...so this is another PR...without admitting the issue. I don't know github's infrastructure, but they have a single point of failure? Last I know, every place these days have backup power especially a datacenter...so those were not working either? My point is that it's much better to be upfront sometimes. In fact github didn't have to say anything about the whole thing since everyone forgot already...
- Animats 11y ago"We identified the hardware issue resulting in servers being unable to view their own drives after power-cycling as a known firmware issue that we are updating across our fleet." Tell us which vendor shipped that firmware, so everyone else can stop buying from them.
- gruez 11y agoI'm guessing they didn't disclose the vendor because they didn't want to be sued for defamation.
- Animats 11y agoTruth is an absolute defense to libel in the US.
- mikeash 11y agoIt doesn't stop you from getting sued, though, it merely stops you from losing. It's pretty reasonable to want to avoid a lawsuit you're absolutely certain you could win.
- Animats 11y agoVendors very seldom sue customers for publicly saying their product is defective. The negative publicity tends to backfire. Legal action can backfire even worse. If the vendor claims the product isn't defective, they have to prove that in court to win a libel action. That means discovery and examination of the company's internal documents and the complaints of other customers, all on the record.
- theptip 11y agoAnd/or they want to maintain a working relationship with said vendor. Going nuclear is a good way of getting _exactly_ the minimum level of service that your SLA specifies.
- osoti 11y agoBody Parts Humans have that are Now Useless http://goo.gl/qahHUj http://goo.gl/qahHUj
- gaius 11y agoYou can very clearly see two kinds of people posting on this thread: those who have actually dealt with failures of complex distributed systems, and those who think it's easy.
- timiblossom 11y agoIf you use Redis, you should try out Dynomite at http://github.com/Netflix/Dynomite http://github.com/Netflix/Dynomite. It can provide HA for Redis servers