9 ms·
Systemd: Enable indefinite service restarts
- 3abiton 3y ago[flagged]
- tadfisher 3y agoThis must be a different philosophy. When I see something like this happening, I investigate to find out why the service is failing to start, which usually uncovers some dependency that can be encoded in the service unit, or some bug in the service.
- chpatrick 3y agoIf your server has a bug that makes it crash every two hours you still want it up the rest of the time until you fix it.
- tekla 3y agoOf course you understand you can do both, like I do.
- zhengyi13 3y agoI think the author's specified use case is to address transient conditions that drive failures. When the given (transient) condition goes away (either passively, or because somebody fixed something), then the service comes back without anyone needing to remember to restart the (now dead) service. By way of example, I've run apps that would refuse to come up fully if they couldn't hit the DB at startup. Alternatively, they might also die if their DB connection went away. App lives on one server; DB lives on another. It'd be awfully nice in that case to be able to fix the DB, and have the app service come back automatically.
- ot 3y agoImagine you use systemd to manage daemons in a large distributed system. Crashes could be caused by a failure in a dependency. Once you fix the dependency, you want all your systems to recover as quickly as possible, you don't want to go through each one of them to manually restart things. This doesn't mean that you don't investigate, it just means that you have an additional guarantee that the system can automatically eventually recover. If you set a limit on number or time or restart, what's a reasonable limit? That will be context dependent, and as soon as it's more than a few minutes, it may as well be infinite.
- mise_en_place 3y agoThat's exactly why systemd should blindly attempt to restart the service infinitely. Seperation of concerns. An init system should simply start and monitor services. That is what an init system is meant to do. The fact that systemd is overengineered and tries to do multiple things causes headaches for a lot of us. Busybox-init is one of the best alternatives, I would use that everywhere if I could.
- vidarh 3y agoIt's trivial to make systemd do that if that is what you want, but there are also plenty of cases when that is not what you want and you then end up trying to write crash-proof startup scripts to provide backoff instead of just changing a flag in a unit file. (And if you want a dumb unit system, there are plenty of options which will run just fine under systemd as a single unit so you never have to actually use systemd for your own services even if you're forced to use systemd for the overall system for whatever reason)
- BarbaryCoast 3y ago...and now you know why I don't run systemd. I believe their thought process is: what would Windows do? This is an example. For instance, the desktop shell still crashes often. In the old days, this would lock up the keyboard and mouse, and you'd have to power cycle. But MS "fixed" it by simply adding infinite restarts to the system. Now we have systemd. When something crashes, there's no need to fix the bug, just restart it. My favorite new misfeature is PulseAudio. These geniuses actually built code for a multi-user, multi-tasking OS...which will only run for ONE user, and then only if that user is logged in. So forget running cron jobs, and sounding an alert if something needs attention. This is all code produced by FreeDesktop[.]org. Thanks to them, your industrial strength, mission-critical server OS is now only suitable for single-user desktop systems.
- akira2501 3y agoI've always preferred daemontools and runit's ideology here. If a service dies, wait one second, then try starting it. Do this forever. The last thing I need is emergent behavior out of my service manager.
- freedomben 3y agoSystemd can do that exactly that. it just doesn't do that by default. But if that's what you want, it's trivial
- akira2501 3y agoIs it possible to do this system wide? Or do I have to do it for each individual service? It may be a trivial amount of work but if the configuration is fragile, I've gained nothing.
- izacus 3y agoIt's literally described in the article.
- akira2501 3y agoThey're literally rhetorical questions. So, yes, you can set /defaults/, but you can't /force/ the configuration globally. Which means you still need to examine every single configuration file to understand the behavior of your system as a whole. Hence.. why I called it a fragile mechanism.
- MadnessASAP 3y agoIt's trivial to grep all system service files for the relevant lines to see if they override the default. It is also easy to override that override to set it to whatever you like of so desired.
- 3y ago
- mise_en_place 3y agoI’ve been bitten by the restart limit many times. Our application server (backend) was crash looping, newest build fixed the crash, but systemd refused to restart the service due to the limit. A subtle but very annoying default behavior.
- dijit 3y agoare you saying systemd was refusing to restart after manual intervention?
- mise_en_place 3y agoCorrect, because the startup limit had been reached: `service start request repeated too quickly, refusing to start`.
- dijit 3y agoThats terrifying, systemd shouldn't pretend to be smarter than manual intervention. That violates everything I ever enjoyed linux for, I left Windows because it thought it knew better than me.
- tick_tock_tick 3y agoIt doesn't the user just fucked up. You can always run systemctl reset-failed whatever.service
- MadnessASAP 3y agoIt's not terrifying, it's mildly annoying. It's also fixed with 'systemctl reset-failed'. SysD doesn't know if 'systemctl start' was emitted by the operator or by a badly running script.
- cookingmyserver 3y agoI don't think it should matter if 'systemctl start' was issued by an operator or an external script, it should try to start no matter what. SysD itself should use a different start command or flag that is subject to the limit when trying to restart after it detects a failure to start.
- o11c 3y agoIt would be nice if `RestartSec` weren't constant. Then you could have the default be 100ms for one-time blips, but (after a burst of failures) fall back gradually to 10s to avoid spinning during longer outages. That said, beware of failure chains causing the interval to add up. AFAIK there's no way to have the kernel notify you of when a different process starts listening on a port.
- dijit 3y ago> AFAIK there's no way to have the kernel notify you of when a different process starts listening on a port. You can use mandatory access control for this. AppArmour or SELinux are examples. Unfortunately they are hard, not sexy and sysadmins (people who tend to do not sexy hard things) are a dead/dying breed
- nomel 3y ago> AFAIK there's no way to have the kernel notify you of when a different process starts listening on a port. Would the ExecCondition be appropriate here, minimally, with a script that runs `lsof -nP -iTCP:${yourport} -sTCP:LISTEN`?
- o11c 3y agoI'm talking: once your process has started, how do you wait for a process you depend on? Obviously if systemd opens the port for you it's easy enough (in this case, even across machines), but otherwise you have to do a sleep loop. And I'm not sure how dependency restarts work in this case. ExecCondition must moves the spin to systemd, and has more overhead than doing in your own process. There's no point in gratuitously restarting after all.
- saint_yossarian 3y agoThere's `RestartSteps` and `RestartMaxDelaySec` for that, see the manpage `systemd.service`.
- o11c 3y agoAh, not in the man page on my system. Available since systemd 254, released July 2023 (only 1 release since then). Huh, has release rate severely slowed down?
- halyconWays 3y agoSeems reasonable if the service is failing due to a transient network issue, which takes many minutes to resolve.
- ElectricSpoon 3y ago> I would guess the developers wanted to prevent laptops running out of battery too quickly And I would guess sysadmins also don't like their logging facilities filling the disks just because a service is stuck in a start loop. There are many reasons to think a service failing to start multiple times in a row won't start. Misconfiguration is probably the most frequent reason for that.
- twic 3y agoExactly. If a service crashes within a second ten times in a row, it's not going to come up cleanly an eleventh time. The right thing to do is stay down, and let monitoring get the attention of a human operator who can figure out what the problem is. Continually rebooting is just going to fill up logs, spam other services, and generally make trouble. I'm sure there are exceptions to this. For those, set Restart=always. But it's an absolutely terrible default.
- BenjiWiebe 3y agoIt might actually, if a network connection is temporarily down.
- rendaw 3y agoOr a disk not attached yet. Or another service it depends on being slow to finish starting up.
- bravetraveler 3y agoSo, you two know how systemd gets heat for doing too much, right? This is one of those things. The 'After=' and 'Requires=' directives address this. Depends on a mount? Point those directives at a '.mount' unit. Depends on networking, perhaps a specific NIC? Point those directives at 'systemd-networkd-wait-online@$REQUIRED_NIC.service' Point being: declare these things, don't wait for entropy to eventually become stable.
- deathanatos 3y ago> Why does systemd give up by default? > I’m not sure. If I had to speculate, I would guess the developers wanted to prevent laptops running out of battery too quickly because one CPU core is permanently busy just restarting some service that’s crashing in a tight loop. sigh … bounded randomized exponential backoff retry. (exponential: double the maximum time you might wait each iteration. Randomized: the time you want is a random amount, between [0, current maximum] (yes, zero.). Bounded: you stop doubling at a certain point, like 5 minutes, so that we'll never wait longer than 5 minutes; otherwise, at some point you're waiting for ∞s, which I guess is like giving up.) (The concern about logs filling up is a worse one. It won't directly solve this, but a high enough max wait usually slows the rate of log generation enough that it becomes small enough to not matter. Also do your log rotations on size.)
- deleted 3y ago[deleted]
- kaba0 3y agoArguably, this logic should live in another place that monitors the service. Especially that service startup failure is usually not something that gets fixed on its own, like a network connection (where exponential backoff is (in)famous). A bad config file, or a failed disk won’t recover in 10 minutes on its own, so systemd’s default makes sense here, I believe.
- otterley 3y agosystemd is a service monitor. It wouldn't be nearly as useful if it wasn't!
- gizmo686 3y agoFrom the servers perspective, external problems typically do get fixed on their own. It is nice when resolving the primary issue is sufficient to fix the entire system; instead of needing to resolve the primary issue; then fix all the secondary and tertiary issues. At my work, we have a simple philosophy for this. The tester is allowed to (on the test system): toggle servers' power; move around network cables; input bad configuration; etc; in any permutation he wants. So long as at the end of the exersise everything is setup correctly the system should function nominally (potentially after a reasonable delay). There should, of course, be a system level dashboard that notifies someone there is a problem; but that is unrelated to the server internal retry logic.
- franknord23 3y agoI believe this allows you to have cascading restart strategies, similar to what can be done in Erlang/OTP: Only after the StartLimit= has been reached, systemd considers the service as failed. Then services that have Required= set on the failed service will be restarted/marked failed as well. I think you can even have systemd reboot or move the system into a recovery mode (target) if an essential unit does not come up. That way, you can get pretty robust systems that are highly tolerant to failures. (Now after reading `man systemd.unit`, i am not fully sure how exactly restarts are cascaded to requiring units.)
- vidarh 3y agoYou can trigger units explicitly on failure with OnFailure=someservice as well (and since you can parameterize service names, you can have e.g. a single failure@.service that'll do whatever you prefer once a service fails. OnFailure makes it easy to implement more complex restart or notification logic.
- twinpeak 3y agoRecently discovered while making a monitoring script that systemd exposes a few properties that can be used to alert on a service that is continuously failing to start if it's set to restart indefinitely. # Get the number of restarts for a service to see if it exceeds an arbitrary threshold. systemctl show -p NRestarts "${SYSTEMD_UNIT}" | cut -d= -f2 # Get when the service started, to work out how long it's been running, as the restart counter isn't reset once the service does start successfully. systemctl show -p ActiveEnterTimestamp "${SYSTEMD_UNIT}" | cut -d= -f2 # Clear the restart counter if the service has been running for long enough based on the timestamp above systemctl reset-failed "${SYSTEMD_UNIT}"
- PhilipRoman 3y agoI can understand avoiding infinite restarts when there is something clearly wrong with configuration, but I can't figure out why they made the "systemctl restart" command also limited by this. For services which don't support dynamic reloading, restarting them is a substitute for that. This makes "systemctl restart" extremely brittle when used from scripts. Nobody accidentally runs "systemctl restart" too fast, when such a command is issued it is clearly intentional and should be always respected by systemd.
- bravetraveler 3y ago> And then you need to remember to restart the dependent services later, which is easy to forget. You missed the other direction of the relationship. I posted elsewhere in the thread on this, don't rely on entropy. Define your dependencies (well) After=/Requires= are obvious. People forget PartOf=.