12 ms·
Update on 1/28 service outage
- bjacobel 11y agoNot much detail here. A more thorough postmortem would give me more confidence they can recover from another similar issue. Hoping to see one soon.
- Zikes 11y agoI agree that a postmortem would be great, but it's good PR for companies to quickly put out statements like this to admit fault and maintain customer trust.
- anon987 11y agoYep, I think most of these post-mortems from any company are pointless from a technical perspective. It's 4 paragraphs that boils down to "someone did something wrong and we'll make sure it doesn't happen" with zero specifics. There's no point in reading these because there's no technical information. Stuff like this is something you sent to your customer because they want root cause.
- Zikes 11y agoI strongly disagree that these sorts of communications are pointless. In every major service outage I've seen where the company maintained a degree of silence, it's caused major damage to their public relations and consumer trust. I know it doesn't tell you much about exactly what happened, but the truth is they may still be sorting that out and focusing on ensuring it does not happen again. An in-depth post-mortem accompanied by a description of the fix would be great. In the meantime, admitting culpability and apologizing are the ideal essential first steps.
- outworlder 11y agoGive them time.
- frik 11y agoYou can see the cascade effect on their status page graphs: https://status.github.com/ https://status.github.com/
- Loic 11y agoWhat is impressive is that with a website 2h down, they can still announce a 97% availability for the day even so the graph clearly shows the 2h of failures in the day... :-/
- arthurschreiber 11y agoUnless I'm mistaken, 97% of (24 hours) = 23.28 hours.
- mirekrusin 11y agoyes, it went down to 89% or something just after the problem.
- deleted 11y ago[deleted]
- WillAbides 11y agoThe 97% you see on the status page is for the past 24 hours. That doesn't include any of the outage being discussed here.
- ceejayoz 11y agoInteresting that their exception logging didn't get turned back on until this morning, from the looks of things.
- rcthompson 11y agoWell, if exception logger was going off nonstop due to the outage, yet not providing any new information, it would make sense if they disabled it until things had returned to normal.
- deleted 11y ago[deleted]
- Zikes 11y agoI'm starting to think that people should mirror their packages to BitBucket as a rule, and that package managers should round robin/flip a coin between the two, or use whichever is available in case of outages.
- rch 11y agoI'd rather have something like Netflix's Open Connect Appliances, covering all of Github, sitting in each office and a centrally located colo facility.
- Zikes 11y agoI'm not familiar with Open Connect Appliances, but Github as a platform is still a Single Point of Failure at least on some level. Domain or SSL issues, for example. I think that as long as we have options to host packages on other platforms in addition, it should be seriously considered. At the very least, it would encourage a more competitive atmosphere for open source hosting services.
- deleted 11y ago[deleted]
- tommoor 11y agoThis post makes it sound like Github has it's own data centers and power infrastructure structure, this is definitely news to me.. I'd presumed co-lo at best.
- noazark 11y agoThe last news I've heard about it was back in 2009, https://github.com/blog/493-github-is-moving-to-rackspace https://github.com/blog/493-github-is-moving-to-rackspace. But I've also heard that they have some infrastructure on site (clearly not what they were talking about).
- seiji 11y ago"data center" is a confusing term. Very few companies build their "data centers" (apple, google, amazon, NSA, actual 'data center' companies, etc). Most companies rent cage space in a larger data center and call that their "private data center." Smaller companies will rent a few dedicated servers or colo half racks from other resellers.
- brazzledazzle 11y agoUnless it's explicitly stated to be a wholly owned data center I always assume companies are talking about rack space in a bigger DC like supernap.
- smaili 11y agoIt's always scary when a cloud service you rely on goes down but great to see GitHub recover. Well done!
- bhaak 11y ago"Millions of people and businesses depend on GitHub" Well, we shouldn't depend on it so much. I shudder at the thought what an outage of GitHub would mean for our company. This time, we were lucky as it was during the night in Europe. Unfortunately, I don't have the power to test this scenario in our company.
- cookiecaper 11y agoIt shouldn't really have much effect. One of git's major selling points is that it's a DVCS, meaning that everyone has a local copy of the repository. Perhaps some collaboration features will be down for a couple of hours (which I think is a downside to GitHub's decision not to put issues/PR history inside of git), but everyone should still be able to do work, commit to the repo, review history, and so forth. If you have people who do code, they can probably find something to work on for two hours without having the Issues/PR interface, right?
- Domenic_S 11y agoAll sorts of other dependencies go down though. Packages you need for your build aren't there. CI or testing integrations don't happen. Code review is probably not happening. If you track issues in GH you can't see what's next to work on or look up requirements. You're right in that you're (probably) not totally deadlocked. But I can't start to estimate the lost $$ in productivity that comes with a global GH outage because of all that.
- raverbashing 11y agoHave a local repos that mirrors the master one on GitHub periodically Should that fail, start working on the local repos until github is back, then sync back to it
- pc86 11y agoDepending on your definition of "periodically" you may lose almost as much time to syncing back than the outage would have caused without the local mirrors.
- skewart 11y agoAm I the only one who is a little shocked that a power outage could have such a huge effect and bring them down for so long? I'm not an infrastructure guy, and I don't know anything about Github's systems, but aren't data center power outages pretty much exactly the kind of thing you plan for with multi-region failover and whatnot. Is it actually frighteningly easy for kind of to happen despite following best practices? Or is it more likely that there's more to the story than what they're sharing now?
- LinuxBender 11y agoI am not at all surprised. There are 'best practices' and then there is what really happens based on business processes and needs. In reality, even the most cloudy of cloud providers will run into this problem at some point. Folks often come up with ideas of implementing something like Chaos Monkey in their data-center, then realize the actual impact it will have and find it is almost impossible to get the rest of the business to agree to this concept. It isn't as easy at it sounds. I only know of two businesses that have actually implemented Chaos Monkey; one being the company that coined the term. Even regular reboots won't catch these problems and if folks were honest, you would see +1 year up-times on most servers in most places. That is just based on my experiences and I am sure some of you have seen different.
- JetSetWilly 11y agoThe problem is most environments are very heteregenous. I evaluated chaos monkey approach for a big bank, the issue is that netflix has whole data centres full of loads of machines doing pretty much the same thing, streaming and serving. And the worst that can happen is a customer's stream stops and they have to restart it. But in most big companies you have thousands of apps that are all doing very different things. Perhaps a critical app might run on 4 hosts spread across two data centres - you're not going to convince people to have chaos monkey regularly and randomly bringing down these hosts, it would cause real impact and is risky. Yeh in theory it should be able to cope but in reality the scales in most orgs are quite different. That said github sounds a lot more like the netflix end of the scale, doing one specific thing at large scale.
- moondev 11y agoGithub doesn't deploy their services in multiple az's?
- rs999gti 11y agoMaybe they do. But this two hour failure tells me that they have never really tried a hot failover and failback scenario in order to test the resiliency of their site.
- detaro 11y agoOr something happened that didn't happen in the tests. And if they suspected something might be in an inconsistent state, taking some downtime to make sure it comes back up properly clearly is the better option.
- moondev 11y agoHope we get more info about it. Would be very interesting to see how their architecture is setup
- anton_gogolev 11y agoIt's one thing when one temporarily loses access to remote repositories for pushes. Quite bearable, because you can exchange code across your corporate network using patches and whatnot. And it's totally different when you cannot friggin build anything because package managers grab dependencies directly off of GitHub.
- msbarnett 11y agoThis is more an argument for caching or vending dependencies than anything else. If the ability to make builds is critical to your org, making your build process depend on the availability of third-party services over which you have no control is going to end in tears.
- banku_brougham 11y agoThis is it. Production builds have to have dependencies hosted internally, not all over the web.
- saidajigumi 11y agoAgreed. The modern ease of pulling in third-party dependencies, while wonderful in its way, has gotten so easy that even "simple" applications require automated caching infrastructure. E.g. if you just fork your top-level dependencies, you won't pick up any of your recursive dependencies. I suppose we all need package manager and git/VCS aware recursive forking/caching tools now. E.g. works with npm, gem, etc. and recursively forks your entire dependency chain. And to think that I managed that sort thing of entirely by hand some years back. (For C/C++ libs, then, so far more manageable.)
- ibejoeb 11y agoFor those that have been affected by this, what parts of your process were disrupted? I've read, so far: * Build fails due to unreachable dependencies hosted by GitHub * Development process depends on PRs
- free2rhyme214 11y agoChinese DDoS? Somehow I don't buy power going out at a server farm.
- cjbprime 11y ago> Chinese DDoS? Somehow I don't buy power going out at a server farm. You should read more about server farms.
- oxguy3 11y agoWhy not? Things break. Electricity is one of those magical things that's very hard to have insanely good uptime -- frankly, it's incredibly impressive that power outages aren't more common. And why would GitHub not disclose that it was a DDoS? They were very forthcoming when there actually _was_ a Chinese DDoS last April: http://arstechnica.com/security/2015/04/ddos-attacks-that-crippled-github-linked-to-great-firewall-of-china/ http://arstechnica.com/security/2015/04/ddos-attacks-that-cr... And in a DDoS, the service typically becomes slower and slower until it reaches the point where only like one in a hundred requests succeeds. With the GitHub outage, it died fairly instantaneously, and it was completely 100% dead. There was no timeout as the servers tried to respond -- the "no servers are available" error page loaded instantly every time.
- johnhenry 11y agoConsidering the attacks within the past year, I was thinking the same thing. I hate to spread conspiracies without foundation, but I wonder if anyone has seen a assessment on the cyberkinetic capabilities of nations around the world?
- rburhum 11y agoYesterday I was being a bit of an ass to a few people about how "the whole point of using git is so that we can do decentralized code management and why these dependencies were being pulled from our private github if the could be sent point to point yadda yadda yadda". Then they proceeded to go over the list of package managers and dependencies we used and I had to shut up. Even when we host our own Docker Hub and package managers (we do), if you dig far enough, you can find some dependency of a dependency of dependency that relies on GitHub. Brew/npm/build script/whatever. It is crazy how everything has changed so much in the past few years. GitHub went from something that was really nice to have to a core requirement for complex systems that rely heavily on open source.
- viperscape 11y agoThe package system for the rust language actually relies on github, as many found out during outage. I don't know if that will change, probably will with a read copy in a different git service.. but I thought it was interesting because I use github for everything save a few private projects, as I imagine most do. I'm not sure what to think of this, it seems backwards and grossly incompetent, yet here we are using it almost exclusively. It might be smart to decentralize some of this with torrents, if that's possible. Even if it was the read portion of a repository, it seems like something to consider, if it hasn't been already
- rms_returns 11y agoNot just rust language, to the best of my knowledge, even packagist, the php package manager relies heavily on github for sourcing its packages. But I think they have other resources too, apart from github.
- jdminhbg 11y agoRuby's bundler doesn't entirely rely on Github, but pulling from a Github repo is a supported option that many take advantage of.
- tatterdemalion 11y ago
- out_of_protocol 11y agoVarious date/time formats across the world bringing me to the knees. If 1/28 outage was _that_ rough 2/28 would be twice as bad and 28/28 would feel like armageddon maybe?
- beachstartup 11y agoit seriously makes me lol that people are upset, or surprised, that an internet service went down for a couple of hours. a couple of hours! get some perspective please. go for a walk, get a tasty burrito, try a new brand of hot sauce. "why didn't they do X, Y, or Z" the answer in every case is it's extremely expensive, or extremely hard to do, or both. you want a reason, there's the reason. maybe they'll fix it. maybe they won't. next question. make your own backups and redundant systems. "but github is so critical!" -- even more reason to have a backup. bad shit happens in this world. even to good people. prepare or suffer the consequences.
- gavazzy 11y agoWould it be possible for a cross between Git and Torrents? Rather than having a central server to pull/push from, instead the server would provide a list of clients. If the server goes down, the list is still available, and so people who depend on it would be able to communicate.
- ljk 11y agoMaybe I'm ignorant, but why do companies rely on github? Why not just host it in-house? If there's power outage in the office then everything would be down anyways, right?
- danneu 11y agoA rare two hour Github outage isn't enough to make anyone on my team want to start dicking around with internal tools.
- ryanfitz 11y agoI recently read a blog post from Github about them operating their own datacenter http://githubengineering.com/githubs-metal-cloud/ http://githubengineering.com/githubs-metal-cloud/ Im not positive, but it sounds like a fairly recent switch from a cloud provider to their own datacenter. If thats the case, Id expect a number of outages to come in the following months.
- secure 11y agoAFAIK, they never used a cloud provider.
- ryanfitz 11y agoGithub was hosted at rackspace, here is there blogpost about it https://github.com/blog/493-github-is-moving-to-rackspace https://github.com/blog/493-github-is-moving-to-rackspace From their blog posted last month: As we started transitioning hosts and services to our own data center, we quickly realized we'd also need an efficient process for installing and configuring operating systems on this new hardware.
- nickpsecurity 11y agoHere's the only page I could quickly find on Github's architecture for those interested: https://github.com/blog/530-how-we-made-github-fast https://github.com/blog/530-how-we-made-github-fast This looks like a single datacenter. I don't see anything here indicating high availability or other datacenters. You'll usually spot either an outright mention of it or certain components/setups common in it. They might have updated their stuff for redundancy since then. However, if it's same architecture, then the reason for the downtime might be intentional design where only a single datacenter has to go down. Might be fine given how people apparently use the service. It's just good to know that this is the case so users can factor that into how they use the product and have a way of working around the expected downtime if it's critical to do so.
- matt_wulfeck 11y agoWhy is it so hard for us to distribute our dependencies? Hash the package to a sha and put t anywhere on the Internet. Then we just need a service that holds and updates the locations of the hashes and we can fetch them anywhere.