4 ms·
Is this not a solved problem? You have the load balancer stop sending new connections to the host. You send sigterm to the process. You wait until timeout, t
by keypusher 8y ago
Is this not a solved problem? You have the load balancer stop sending new connections to the host. You send sigterm to the process. You wait until timeout, then you sigkill the process. This is how it's done in Kubernetes, ECS, and other systems I've worked with. Trying to engineer the entire lifecycle from within the webserver has a wide array of problems that the author is bending over backwards to end up only partially solving.
- rumcajz 8y agoHere's a scenario: You webserver has an open long-lived WS connection. You stop load balancer. The connections still lives. Then you send SIGTERM to the webserver. The connections still lives. The main coroutine of the webserver catches the SIGTERM and wants to do gracefull shutdown. But the connection in question in running in a separate coroutine. So it has to send a graceful shutdown signal to that coroutine. Etc. You end up with having to deal with the problems described in the article.
- ninkendo 8y agoIf you have a websocket architecture, the client needs to tolerate a severed connection and transparently reconnect, without user-visible impact. Anything else is going to lead to disaster... the server end can't stay up forever. The right thing to do in this case is for the server to also disconnect any websockets when it gets the SIGTERM, in addition to stopping the accept loop and the other things it would normally do. Clients simply have to reconnect (and have enough tolerance in the socket's in-band protocol to recover state, roll back transactions, etc.) It's a scenario that's bound to happen to a certain percentage of your clients anyway, due to all sorts of external factors, and must be handled regardless.
- rumcajz 8y agoCLOSE frame exists in the WS protocol for a reason. It's there so that both sides know that all the messages were delivered before the connection was closed.
- adontz 8y agoWhy coroutine is long lived? I image this as event loop machinery just stops processing events and all sockets are closed [by OS].
- adontz 8y agoActually you do not need even load balancer node or software. SO_REUSEPORT transforms Linux kernel into load balancer. You just run new server alongside old one and ask old one to stop serving requests.
- Twirrim 8y agoOne note of caution: One way to handle this is by having your application stop responding to health checks from the load-balancer. The LB marks the host as dead and stops sending new connections to it. However there are some load-balancers that see a failed health-check and immediately flush all connections to a "failed" host. Be sure you know what your LB's behaviour is, and how that fits in to your graceful shutdown model.