4 ms·
This peculiar behavior (where the Service health check is unaware of the Deployment’s known instances) mirrors Google Compute Engine (where the httpLoadBalancer
by yonran 5y ago
This peculiar behavior (where the Service health check is unaware of the Deployment’s known instances) mirrors Google Compute Engine (where the httpLoadBalancer’s healthCheck is independent of the instanceGroupManager’s known instances). If you want your program to exit as soon as possible rather than waiting for a SIGKILL from the instance group manager, you have to hard-code the health check timeouts like so:
1. After a SIGTERM, the shutdownHook should keep the HTTP server running. Future /@status requests must return an error, but user requests must still succeed!
2. The shutdownHook sleeps for a minimum of the load balancer health check’s checkIntervalSec * unhealthyThreshold + timeoutSec (which by the way must be less than the instanceGroupManager’s health check’s checkIntervalSec * unhealthyThreshold if it uses the same endpoint)
3. Now the load balancer should not be sending new requests. The shutdownHook then waits for any existing requests to drain.
4. After requests drain, the shutdownHook can finally exit gracefully.
It is annoying to have to wait for the health check’s delay (rather than simply draining existing requests as in AWS), but it seems to be necessary for Google-designed load balancers and instance group managers.
- paulfurtado 5y agoWith kubernetes you can at least add a preStop hook that sleeps for 60-120sec if the app is not designed to handle SIGTERM. The pod enters the terminating state just prior to executing the preStop hook.