3 ms·
Github down, no hard drives available, no memory available, thanks AI! Seems like we are headed for Tech Gridlock.
by annoyingnoob 2mo ago
Github down, no hard drives available, no memory available, thanks AI!
Seems like we are headed for Tech Gridlock.
- jdm2212 2mo agoThis stuff is good! This is what a booming economy looks like. There are people out there competing with you for resources because they have cool ideas they want to implement.
- a2ff6eeb0 2mo agoOr at least they asked the AI to come up with cool ideas, which is even more interesting. It's exciting watching the world transition away from humanity being in the driver's seat!
- yoyohello13 2mo agoLayoffs by the 10s of thousands, food prices out of control, people barely able to afford gas. At least some tech bros can launch their 50th B2B SaaS. I didn't realize a booming economy sucked so much.
- jdm2212 2mo agoThe exciting thing about AI is precisely that it'll let software move beyond Yet Another B2B SaaS and into doing useful things in the real world. I regularly ride in driverless cars! That was the stuff of science fiction when I was a kid. If you're worried about food prices, you should be happy that robots will make agriculture less labor-intensive and bring prices down.
- kyleweng 2mo agoprobably worth asking when those prices will come down.
- jdm2212 2mo agoAutomation has been making food steadily cheaper for 300+ years, and continues to do so. Sometimes the price of oil goes up so much that fuel and fertilizer price increases negate the wins from automation, but automation is still generating wins and will only accelerate as robots get smarter.
- yipinwong 2mo agoWhat they can implement is to slowdown the commit rate, rate limt or just queue-up messages not to overburden their downstream service. I don't think GH has any of those, but just keep scaling, but that scaling failed. Just bad architectural decisions from the postmortem. -- It will only get worse due to AIs spawning massive commits, and they don't have unlimited cloud resource. They can scale but not scalable in terms of effort, resources, and $
- jdm2212 2mo agoHow would any of what you're saying help with this? > The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.
- yipinwong 2mo agoload balancer failure? rate limit woudl address concurrency limits? throttle or queue up messages. auto-scale failed cause was misconfiguration policy, which i admit cannot be handled by my suggestions. The cascade? it's downstream service degradation, which I mentione should have had been prevented with queues. One of the jobs that queues/kafka solve is to prevent these downstream outages.
- jdm2212 2mo agoIf your LB is down, you're just kind of screwed. You can't enqueue things if requests aren't getting through at all. Same deal with authn/authz issues, which they also had. If you can't answer the question "is this message allowed to be added to the queue" you can't enqueue stuff. GitHub does use queueing for all kinds of stuff internally, though, because they're not morons.