4 ms·
for long-running daemon/server programs where reliability is a key priority, I have had a lot of success building in "partial" restarts into the design, which
by jstrong 6y ago
for long-running daemon/server programs where reliability is a key priority, I have had a lot of success building in "partial" restarts into the design, which means that periodically the program exits the main event loop and enters it again.
clearly there are circumstances where this is not an option, but for distributed systems that can handle the short-term loss of a given node it works very well. one important thing is to add is some randomized jitter to the periodic restarts to spread out when nodes are offline.
in practice this remedies a wide range of seldom-encountered problems. for instance, I have a program that pulls data from an http apis, and last week its connection to one of the remote servers became hung up somehow, resulting in no new data being pulled from that server. this problem had never happened before in 6 months since the program was deployed. rather than track down the very rare condition that was behind this I just added periodic "partial" restarts to the program, and now it will be able to recover from this if it ever happened again, and many other sorts of problems like slowing increasing memory fragmentation.
for a chat server or other server with client connections, obviously those would need to be maintained across the partial restart, but that seems like it would be pretty easy to do. while there is nothing dramatically different from my approach to what you did with actual, full periodic restarts, the partial restart approach does allow special handling for application-specific issues like that.