3 ms·
Sometimes its not a good idea for the health of a service to be determined by its connected parts (eg databases). For purely situational awareness this is fine.
by coenhyde 7y ago
Sometimes its not a good idea for the health of a service to be determined by its connected parts (eg databases). For purely situational awareness this is fine. But if you use the healthcheck to determine if an instance of an application should be taken out of service you risk cascading failures; turning 1 problem into 10. It's usually better for the application to throw an error if it can't connect to the database. That said I do both approaches depending on the situation.
- derefr 7y agoI mean... that's what circuit breakers are for. If a component of a service is optional to its operation, then it wraps calls to that service in a circuit-breaker and fails requests that ask for that service. And if a component of a service is not optional to the operation of a service, then the failure should cascade to the service's dependent clients, and their dependent clients, and so on, so that there's backpressure all the way back to the origin of the request.
- closeparen 7y agoHealth checks can have several purposes. They're used by the routing control plane to determine inclusion in the load balancer pool. This is already a kind of circuit breaker and is similar to what an application-level circuit breaker would poll. But they're also used by the scheduler to determine whether the instance needs to be restarted. You don't want your thing in a restart loop just because a dependency is down! In fact, if a very widely shared dependency went down, and everyone was checking it in their health check, the scheduler control plane could quickly have a backlog measured in days trying to move all those instances. Our environment now supports distinct answers to those two questions but most service authors don't know about it.
- hinkley 7y agoTwo modes of cascading failure here: Request of Death, and cascading failures. If a request kills a particular server you should let the error flow upstream, otherwise it will just bounce from server to server until it's killed all of them. For the latter, someone related a real-world example of this to me the other day. Say you have a bunch of people managing customers. Every employee has 4 customers, and those take up all of their time. You get a new customer. Instead of hiring a new rep, you give someone a 5th customer to manage. They struggle, and eventually they quit. Now, all of your employees have 5 customers. Sooner or later one of those will also quit, and then it's a race to see who can get out the door fastest. The moral of that story is that all the load balancing in the world is for naught if you haven't done your capacity planning properly. And once the system starts to buckle it may be too late to bring new capacity online (since startup usually consumes more resources).