5 ms·
Runit is amazing. I've used it on several large scale websites with great success. Runit follows the unix philosophy of being stupid simple and doing one thing
by atom_enger 11y ago
Runit is amazing. I've used it on several large scale websites with great success. Runit follows the unix philosophy of being stupid simple and doing one thing incredibly well. If you're starting up a new project, consider using Runit.
- paulsmith 11y agoNot only is it simple and does the do-the-one-thing-well thing well, which is true, it's also _correct_. It's hard to overstate how valuable runit is in production because of that. It does the right thing with regard to clearing the environment, detaching from controlling terminal, logging, and many other subtle aspects of operating a service. I never worry about runit.
- dap 11y ago> Not only is it simple and does the do-the-one-thing-well thing well, which is true, it's also _correct_. How does it address the "who-watches-the-watcher" problem typically associated with service restarting?
- ori_b 11y agoMake it simple and obviously correct, and don't crash.
- dap 11y agoJust don't make any mistakes? Was that a joke? There are many reasons a process can die that are outside of its control, including signals from outside the process, handled (but uncorrectable) memory errors, and the OOM killer (on Linux). Besides that, it seems like a major design shortcoming if fatal errors in any particular program (however critical and however simple that program may be) can be unrecoverable for the whole system. It's definitely possible to solve this problem rigorously and completely, though I don't know of a way to do it without support from the kernel. On illumos systems, the service restarter ("svc.startd") provides a complex restart policy for user-defined services. I believe the restarter itself is restarted blindly by init, and init is restarted blindly by the kernel. If the kernel dies, the whole system is rebooted. In this way, if any software component in the chain of restarters fails, the system still converges to the correct state.
- ori_b 11y agoWhat if init doesn't exit, and just hangs? What if it just goes crazy and starts erronously restarting your processes? There are more failure modes than simply crashing. At some point, you just have to assume that some critical components are working correctly. Adding complexity just makes it harder to reason about it, or, depending on how paranoid you are, prove it.
- dap 11y agoThat's a slippery slope argument: because we can't solve the halting problem or verify program correctness, we shouldn't try to handle crashes, either? Agreed on minimizing complexity. The only part of the chain I described that's very complex is svc.startd, and that's largely to support rich configuration. Also don't mistake my position for saying that quality isn't important. Rather, just that perfection is not a reasonable constraint.
- vezzy-fnord 11y agoAt least on Linux, PID1's death is an instant kernel panic. As such, it's wise to keep any service management logic out of it. If the svscan process dies, then your system is still chugging along and you can intervene to restore the supervision tree (otherwise svscan inspects supervise processes at a regular 5s interval). If you have some really critical process, then you could integrate a checkpointer into the run script chain so that you can just pick off from the last image of the process state with minimal interruption.
- ori_b 11y ago> Also don't mistake my position for saying that quality isn't important. Rather, just that perfection is not a reasonable constraint. At some point, for some component or set of components, perfection is your only choice, regardless of the rest of your design. At least when you consider a single node with a single point of failure; this is less true for a distributed system where you have redundancy. At some point, you have to assume that either init is perfect, or that the code in the kernel to detect init failures is perfect, or that the watchdog monitoring the kernel is perfect, or whatever other layering you choose to put in place is perfect. In a system with a finite number of components, there is always going to be a point at which you just say "this bit is going to have to be correct, and there's no other way around it".
- freshhawk 11y agoBack when I changed from supervisord to runit my life was substantially improved. That correctness means way fewer emergency maintenance ops issues in production.
- general_failure 11y agoWhat is wrong with supervisord?
- freshhawk 11y agoFirst of all and mostly, we could never exactly figure it out what was going wrong. This was a few years ago and I don't remember all the details. We just had occasional issues with stopping/starting and especially restarting processes when pushing a new version out or when a process crashed. I do remember it could occasionally report a successful restart and still leave the old process(es) running. All my developers became very familiar with supervisord, and it was number 2 on the troubleshooting list (1. Did we introduce a bug in a recent commit? 2. Did supervisord do something weird again). After we switched to runit, only devs that touched ops knew about it at all. And we all forgot it was there. That's what you want in an ops tool.
- bnolsen 11y agoand additionally its totally understandable. its not scared of basic shell scripts and leveraging the wealth the posix environment provides.