3 ms·
> We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure. And from the RCA [1]: > The
by dcrazy 2mo ago
> We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure.
And from the RCA [1]:
> The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic.
[1]: https://www.githubstatus.com/incidents/zkxwbgr0cnmx https://www.githubstatus.com/incidents/zkxwbgr0cnmx
- ivraatiems 2mo ago"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more. It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage, unless they have no reporting or their tooling is completely incompetent. If things can fall apart so easily, throwing more capacity at the problem won't fix it.
- dcrazy 2mo agoYou’re torturing your own logic to make Azure the villain here. And it also sounds like you lack experience with capacity exhaustion. Things fail slowly, then suddenly.
- ivraatiems 2mo ago[flagged]
- dcrazy 2mo agoBaseless accusation made from a position of zero information.
- ivraatiems 2mo agoOpinion based on stated facts. Please share the information you have which contradicts the conclusions I have drawn from Github's statement. (And we know they're liars. They report very few of the actual incidents they have; see for example https://mrshu.github.io/github-statuses/ https://mrshu.github.io/github-statuses/)
- dcrazy 2mo agoI don’t owe you anything, much less a separately sourced counterargument. You openly admit that your opinion is not based not on GitHub’s proffered statements but on your self-admitted assumption that GitHub is actively lying in an attempt to cover up an Azure-related root cause.
- cyberax 2mo agoThis absolutely can happen in large systems. If some part of the system is at capacity, then slightly increasing the load can cause it to fall behind and start accumulating a backlog. These backlogs can cause clients to make more retries, exacerbating the problem. Potentially further cascading through the system.
- Kinrany 2mo agoI believe their point is that "system is at capacity" is something they ought to start fixing before the capacity is exceeded
- dcrazy 2mo agoBut then people like OP will claim that the capacity concerns are a lie manufactured to support an unjustified move to Azure.
- cyberax 2mo agoSure. But you might not even be realizing that something is just at the cusp if the load is spiky enough. The art of large system design is to identify and avoid these kinds of chokepoints. And when something happens, propagate the "backpressure" up the stack to avoid queuing. AWS got a fair share of similar outages, so the newer SDKs now try to not exacerbate these kinds of issues: https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html https://docs.aws.amazon.com/sdkref/latest/guide/feature-retr... The original AWS EBS outage is probably the canonical example: https://aws.amazon.com/message/65648/ https://aws.amazon.com/message/65648/
- rcleveng 2mo agoThere's always a cliff, this part is fine. You sometimes know the cliff but often do not.
- stackghost 2mo ago>"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more. Look, I hate Microslop as much as anyone but you'd have to purposely misinterpret TFA in order to arrive at this interpretation. C'mon.
- maccard 2mo agoI’m firmly in the camp of “something stinks at GitHub” but > It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage In my experience, there are hard thresholds that get passed that expose hidden bottlenecks like this. A previous system I worked on we had absolutely loads of headroom by all of our measured metrics, but one day we filled a cache because the value hadn’t been tweaked in recent memory. Plenty of space on disk and in memory, but all of a sudden we went from a very high cache hit rate to a very low cache hit rate, and everything ground to a halt.