20 ms·
The GitHub Availability Report
- realchucknorris 6y agofacebook can learn something from github
- dmpetrov 6y agoExternal production systems depend on GitHub. For FB - it is fine to fail from time to time. Users will be even more productive :)
- kevsim 6y agoExternal production systems unfortunately depend on FB too as we've seen with all the iOS apps crashing due to issues with FB's iOS SDK.
- dx034 6y agoEven more important, Github lives off fees paid by companies. They might switch to Gitlab or other competitors if availability remains an issue. Facebook lives off ads, as long as people visit FB, companies won't really take their ads somewhere else.
- markwaldron 6y agoGood on GitHub for being transparent. That said, I hope they return to being a more reliable platform soon
- RcouF1uZ4gsC 6y agoA couple of observations: It seems that MySQL seems somehow connected to all the outages, especially the unexpected crashes. >GitHub’s monitoring systems currently alert when tables hit 70% of the primary key size used. We are now extending our test frameworks to include a linter in place for int / bigint foreign key mismatches. In 2020, should we have Integer(32-bit) primary keys anymore? I think at this time, everyone should just go with BigInt or UUID for primary keys/foreign keys, and basically not have running out of key space be an issue you have to worry about.
- minxomat 6y agoI always use UUID PRIMARY KEY DEFAULT uuid_generate_v1mc() In Postgres. It will give you UUIDs and the larger keyspace, without being excessively random (they're basically almost sequential). I do this always, even if the table has other int columns that look like friendly values that could be used as PKs. Some time in the future they will ruin your day.
- fulafel 6y agoFor others curious about this uuid type: "This function generates a version 1 UUID. This involves the MAC address of the computer and a time stamp. Note that UUIDs of this kind reveal the identity of the computer that created the identifier and the time at which it did so, which might make it unsuitable for certain security-sensitive applications." I guess that's not immune to problems either (i can imagine problems with both mac address and time).
- viraptor 6y agoThis is not the right function. V1mc randomises the MAC. Also it's only one of the possibilities. There are many ways if generating uuids depending on your situation.
- deleted 6y ago[deleted]
- abhishekjha 6y agoIsn't UUID discouraged for a PRIMARY KEY pertaining indexes? What would the B+Tree look like for over 2 billions UUIDS where there is no order within? Also, the cache locality problem.
- tasogare 6y agoThis is why GP wrote "It will give you UUIDs and the larger keyspace, without being excessively random (they're basically almost sequential)". The main point for not screwing up the clustered index is the "almost sequential" part.
- maallooc 6y agoI wonder whether these problems are caused by bureaucratic and incompetent developers, poor SQL engine or both.
- dragonwriter 6y agoBureaucratic and incompetent engineering management is more likely than bureaucratic and incompetent developers.
- aaomidi 6y agoThey're the same thing.
- aprdm 6y agoWhy so harsh ? If there’s one reality about software serving so many people with so many developers is that there will be bugs and blind spots.
- KingOfCoders 6y agoThree out of four times MySQL was involved.
- mattip 6y agoIf you are running an app that is basically a UI layer over a database, where would you expect significant failures to occur? The users are already debugging possible problems with git locally by simply using it, so UI-database interactions are where all the action happens on products like github.
- KingOfCoders 6y agoWhen I was CTO in some companies we had very different problems, with DBs being only one of them (and thankfully not four hour long outages in two months).
- aneutron 6y agoI might be wrong, but you probably didn't handle as much traffic as Github does. Bugs and outages probability is a asymptotic to normalized traffic. At least intuitively.
- KingOfCoders 6y agoNo we didn't luckily but we also didn't have that much money. Biggest problem was 50k+ people trying to reserve a limited amount of stuff at the same time and 5k+ logins/second (at first we made the mistake to write last_login to MySQL ;-)
- d0100 6y agoHow did you go about saving login/device data? We are going to do this soon in our app and I'm looking for good use cases and solutions
- KingOfCoders 6y ago
- aetherspawn 6y agoYes, a lot of down time for sure, but each of these is quite an unexpected edge-case. I wouldn’t think that many of these issues would be reoccurring as they have added regression tests and process in place for each.
- rsa25519 6y agoYep. And several of them seem like issues that would be very difficult to reproduce outside of production (e.g. overflowing primary key index), so it makes sense that they were not caught earlier
- dx034 6y agoIt's actually an error I've seen multiple times in the past, and I'm not even working with databases full time. I do think it's surprising that no one had thought of implementing at least checks for these conditions.
- capableweb 6y agoYeah, I think most people who touch/interact with backend/database code, even if not actually backend/DB developers, have a reaction nowadays to seeing auto-incremental IDs because it's famously hard to scale to any distributed architecture + introduces issues when you hit the limit. Projects started in the last three years that I've collaborated/been part of have all ditched the auto incremental IDs.
- baq 6y agoAuto incrementing integer ids have some very desirable properties though. Not sure what you replaced them with.
- dijit 6y agoSomething I see more and more is a primary key based on guid/uuid; I'm not fully aware of the merits of either approach, so I'm just putting the info I have at hand out there.
- TekMol 6y agoI don't understand the first one: a shared database table’s auto-incrementing ID column exceeded the size that can be represented by the MySQL Integer type Followed by: GitHub’s monitoring systems currently alert when tables hit 70% of the primary key size So why was there no alert?
- romanhn 6y agoThe latter looks to be the solution put in place to deal with the former. That seems to be the pattern with all of these - first list the issue, then the last paragraph describes the remediation.
- Doxin 6y agoMore importantly, why is their ID type small enough to where this is even a possibility?
- hvidgaard 6y agoThe "default" is a 32 bit int, large enough for almost 4.3 billion records. That is enough for the vast majority of tables, and going to a 64 bit int has some performance implications, which at GitHubs scale is absolutely something they need to take into consideration, just as they should have with the table size.
- dx034 6y agoI've set up systems where I chose 32bit int because I could've never imagined to need more than 4 billion values. AT some points I did (related to deletes) but changing the datatype when you have >1bn records can be next to impossible for 24/7 operations, especially for primary keys.
- rsa25519 6y agoMy interpretation is that "currently" is referring to right now, after that incident. I assume they meant to suggest that the monitoring system was put in place as a result of the incident. Or, I'm wrong. Maybe there wasn't an alert, or maybe the alert was ignored.
- shenli3514 6y agoI'm the VP of Engineering from PingCAP. As a devoted customer and a big fan of GitHub, it’s devastating to see service interruptions and our business impacted by these problems; as the team behind TiDB, we (PingCAP) believe there is something we can do to help solve the database high availability and scalability problem. We would like to propose the database team in GitHub to consider the TiDB platform. TiDB and TiKV can work as a scale-out MySQL and it has been battle-tested in all kinds of hyper-scale scenarios. Additionally, just in case you didn’t do this, we also recommend GitHub considering the Chaos-Mesh project (https://github.com/pingcap/chaos-mesh https://github.com/pingcap/chaos-mesh) to do Chaos Engineering. Chaos Mesh was our internal chaos engineering platform and we open-sourced it on Dec. 31, 2019. It can be used for simulating different kinds of failures including network partition, flaky disk, node outage. You can easily use it to simulate failures in a test environment and confirm that the high availability works as expected.
- ithkuil 6y agoCan it simulate the overflow of a 32-bit autoincrementjng primary key? Should it?
- shenli3514 6y agoDo you mean Chaos-Mesh or TiDB? For Chaos-Mesh, it is not a suitable tool for deterministic unit test case (like inserting an overflow value and expect to get an error). For TiDB, it is easy to use `Alter Table` statement to enlarge the length of a integer column with zero impact on the online business.
- capableweb 6y agoNot sure if you're purposefully missing the question here. ithkuil is not asking if you can change the length of the integer column but rather if you somehow would have prevented the issue GitHub saw here with hitting the limit. So something like automatically updating the length, or giving the developers warnings, or alerts of some sort. My guess would be no, as I haven't seen this in the wild. And if the answer is indeed "no", your comment seems again to plug something that wouldn't actually solve the problem GitHub saw here.
- stunt 6y agoEverything is easy and obvious in retrospect. If you wonder how they missed something as simple as PK size. As part of maturity model, every product should have periodic milestones base on their scale where engineers should reassessment their infrastructure and the choices they made earlier. But, you can always overlook things unless you have a checklist for everything.
- hoseja 6y agoThe last one is really funny to me in a slapstick way and anti-flapping sounds like a really appropriate term.
- jsnell 6y ago> We strive to engineer systems that are highly available and fault-tolerant and we expect that most of these monthly updates will recap periods of time where GitHub was >99% available. So that's them striving for two nines most of the time, which seems to be a fancy way of saying one nine. Be careful you're not setting the bar too high!
- genidoi 6y agoThat number stands out more than anything else in that report. 99% availability effectively becomes the weakest link in the chain. Would they still have pushed this report out if they phrased it as "...where GitHub was only down for an hour for every 99 hours of uptime"
- m0xte 6y agoAlso this is like South Africa Post Office. 99.9% of packages delivered! = you lost 0.01% of them which totalled millions.
- SahAssar 6y agoI always heard that the number of nines was after the point, so that'd be zero nines. Two nines would be 99.99%
- lokedhs 6y agoThat's wrong. Five-9's for example, is 99.999% uptime, or a bit over 5 minutes downtime per year.
- quietbritishjim 6y agoMaybe you're thinkinng of the actual proportion value. 99% can more simply be written as 0.99, which is two nines after the true decimal point.
- kristianc 6y agoThat doesn’t seem to be what it’s saying. Wording is unclear, but it seems to be saying that _for the incidents recounted in these updates_, most of the time the GitHub service itself will be 99% available. That’s not the same as saying they only strive to make GitHub 99% available at all times.
- rhy_bee 6y agoI'm curious about how Github generates 32 bit IDs in a distributed system.
- deleted 6y ago[deleted]
- nickjj 6y ago> GitHub’s monitoring systems currently alert when tables hit 70% of the primary key size This is interesting. I wonder if they query the key size every N seconds as part of their monitor, or if they report the key size to the monitor on write.
- t3rabytes 6y agoWe have a similar setup that just queries for the auto increment value every X period of time, creates next to zero additional load to do so.
- nickjj 6y agoThanks. Yeah that seems sane. The odds of you exceeding your monitoring threshold in 1 interval seems close to impossible too. Such as going from 70% to 100% in 30 seconds or whatever your monitor interval is set to.
- poorman 6y agoWho doesn’t think to use bigint for their auto incrementing IDs in Rails? This seems like it should be a non-negotiable for a company of scale such as GitHub.
- kellenmurphy 6y agoIn other words, so many outages are happening that we're only going to write about it once a month.
- ksec 6y agoI am wondering if Github ever consider switching to Postgre?
- daiyanze 6y agoInteresting...