3 ms·
Circuit breaking: also called “kill switch”. It’s having a way to shut down some feature if it becomes problematic. “Bulkheading” is making sure that failures
by cnasc 6y ago
Circuit breaking: also called “kill switch”. It’s having a way to shut down some feature if it becomes problematic.
“Bulkheading” is making sure that failures in one area don’t cascade into causing failures elsewhere analogous to how bulkheads on a ship prevent a single breach from sinking the whole ship.
- ignoramous 6y agoAs far as I know, the original source for these patterns is Michael Nygard's book Release It: https://www.amazon.com/gp/product/0978739213 https://www.amazon.com/gp/product/0978739213
- sciurus 6y agoA fantastic book! The first edition is a bit dated in places, but there's a second edition now: https://www.amazon.com/Release-Design-Deploy-Production-Ready-Software-dp-1680502395/dp/1680502395/ https://www.amazon.com/Release-Design-Deploy-Production-Read...
- dtech 6y agoThe difference between a circuit breaker and a kill switch is that a circuit breaker, like a real electrical one, automatically trips and stops requests going through for some time after a certain error threshold to the remote server has passes.
- hayst4ck 6y agoThe context is somewhat important. While 508s were being sent, there was also a significant number of 503s (IIRC). A 503 is an absolute hallmark of something reaching max utilization. Sometimes it's a bad code push that results in memory bloat and then swap or significantly increased request handling time on a poorly threaded server (think a code push that makes a blocking linear request in a loop), but the vast majority of time an upstream dependency (specifically a data store) has been overloaded. So for whatever reason a data store gets slower. What happens upstream when this happens? The number of incoming requests is constant, but the time each individual thread spends attempting to talk to the data store is constant (or worsening). This means each request blocks longer (resulting in potential thread starvation) or there are more concurrent requests (load) to the data store. This creates a feedback loop of doom: as a data store slows down, its load (the number of requests it's handling at once) increases until complete failure. The only way to stop this behavior is by "failing fast." This is how a circuit breaker works. When the data store starts responding slowly, it’s important not to hammer it with even more load, so your client watches the number of load related failures (or response times) and automatically fails requests immediately without sending it to the data store (circuit breaker). This allows the data store to become unloaded and get out of the doomed feedback loop.