4 ms·
Yes, a lot of down time for sure, but each of these is quite an unexpected edge-case. I wouldn’t think that many of these issues would be reoccurring as they ha
by aetherspawn 6y ago
Yes, a lot of down time for sure, but each of these is quite an unexpected edge-case. I wouldn’t think that many of these issues would be reoccurring as they have added regression tests and process in place for each.
- rsa25519 6y agoYep. And several of them seem like issues that would be very difficult to reproduce outside of production (e.g. overflowing primary key index), so it makes sense that they were not caught earlier
- dx034 6y agoIt's actually an error I've seen multiple times in the past, and I'm not even working with databases full time. I do think it's surprising that no one had thought of implementing at least checks for these conditions.
- capableweb 6y agoYeah, I think most people who touch/interact with backend/database code, even if not actually backend/DB developers, have a reaction nowadays to seeing auto-incremental IDs because it's famously hard to scale to any distributed architecture + introduces issues when you hit the limit. Projects started in the last three years that I've collaborated/been part of have all ditched the auto incremental IDs.
- baq 6y agoAuto incrementing integer ids have some very desirable properties though. Not sure what you replaced them with.
- dijit 6y agoSomething I see more and more is a primary key based on guid/uuid; I'm not fully aware of the merits of either approach, so I'm just putting the info I have at hand out there.
- deleted 6y ago[deleted]
- petters 6y agoIt seems that they still don't know what caused the June 29 outage, though.
- dijit 6y agoI don't wish to discredit the work that's been done by github here; nor do I think it's reasonable to say they're negligent. However, at least two of these cases are actually something that I test for as part of the practise of systems administration at large high throughput companies. Transaction ID wrap-around (and auto-increment capacity) are known-knowns in database administration; The CPU starvation and flap detection mechanisms are also known-knowns in systems administration. The premise that they came out of the left field is disengenuous at best; unless I have "special experiences", which is possible I suppose as I was responsible for 1% of all web-traffic at one point in my life... but github should be much higher than that, I am surprised that they haven't tested the first principle assumptions, and controlled for them. Similar to how a developer sprinkles code with asserts to ensure things that are impossible remain impossible; a sysadmin doesn't take for granted that a service will 'be magic'.
- lathiat 6y agoI'm pretty sure "responsible for 1% of all web-traffic" puts you in the "special experiences" category :)
- dijit 6y agoI guess then, the question is: Why does nobody on the github staff have a similar experiences to me, when they (assumingly) operate at even higher scales? Or, if they do, why were they not in on the design meetings? Or, if they were, why were they ignored?