4 ms·
"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more. It is unbelivable that aload of 2.8b
by ivraatiems 1mo ago
"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more.
It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage, unless they have no reporting or their tooling is completely incompetent. If things can fall apart so easily, throwing more capacity at the problem won't fix it.
- dcrazy 1mo agoYou’re torturing your own logic to make Azure the villain here. And it also sounds like you lack experience with capacity exhaustion. Things fail slowly, then suddenly.
- ivraatiems 1mo ago[flagged]
- dcrazy 1mo agoBaseless accusation made from a position of zero information.
- ivraatiems 1mo agoOpinion based on stated facts. Please share the information you have which contradicts the conclusions I have drawn from Github's statement. (And we know they're liars. They report very few of the actual incidents they have; see for example https://mrshu.github.io/github-statuses/ https://mrshu.github.io/github-statuses/)
- dcrazy 1mo agoI don’t owe you anything, much less a separately sourced counterargument. You openly admit that your opinion is not based not on GitHub’s proffered statements but on your self-admitted assumption that GitHub is actively lying in an attempt to cover up an Azure-related root cause.
- cyberax 1mo agoThis absolutely can happen in large systems. If some part of the system is at capacity, then slightly increasing the load can cause it to fall behind and start accumulating a backlog. These backlogs can cause clients to make more retries, exacerbating the problem. Potentially further cascading through the system.
- Kinrany 1mo agoI believe their point is that "system is at capacity" is something they ought to start fixing before the capacity is exceeded
- dcrazy 1mo agoBut then people like OP will claim that the capacity concerns are a lie manufactured to support an unjustified move to Azure.
- cyberax 1mo agoSure. But you might not even be realizing that something is just at the cusp if the load is spiky enough. The art of large system design is to identify and avoid these kinds of chokepoints. And when something happens, propagate the "backpressure" up the stack to avoid queuing. AWS got a fair share of similar outages, so the newer SDKs now try to not exacerbate these kinds of issues: https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html https://docs.aws.amazon.com/sdkref/latest/guide/feature-retr... The original AWS EBS outage is probably the canonical example: https://aws.amazon.com/message/65648/ https://aws.amazon.com/message/65648/
- rcleveng 1mo agoThere's always a cliff, this part is fine. You sometimes know the cliff but often do not.
- stackghost 1mo ago>"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more. Look, I hate Microslop as much as anyone but you'd have to purposely misinterpret TFA in order to arrive at this interpretation. C'mon.
- maccard 1mo agoI’m firmly in the camp of “something stinks at GitHub” but > It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage In my experience, there are hard thresholds that get passed that expose hidden bottlenecks like this. A previous system I worked on we had absolutely loads of headroom by all of our measured metrics, but one day we filled a cache because the value hadn’t been tweaked in recent memory. Plenty of space on disk and in memory, but all of a sudden we went from a very high cache hit rate to a very low cache hit rate, and everything ground to a halt.