4 ms·
> Have mandatory reboots, never be afraid of reboots. Yes! It is not a badge of honor to have a server that's been online for 365+ days -- it just means you ha
by MiscIdeaMaker99 5y ago
> Have mandatory reboots, never be afraid of reboots.
Yes! It is not a badge of honor to have a server that's been online for 365+ days -- it just means you haven't applied any kernel security updates for a long, long time. :-p
- toast0 5y ago> Yes! It is not a badge of honor to have a server that's been online for 365+ days The badges come at 1000 days and every 500 after that. If you have a minimal kernel, and a limited software set, many kernel security bugs found won't apply to your system and you don't need to reboot for them. Not that you need to be scared of reboots, and you should probably schedule some reboots so you can be confident your systems can reboot, but you also don't need to be scared of uptime or stability, either.
- addingnumbers 5y ago1000 days of uptime is 1000 days of configuration drift since the last time you confirmed your services still boot properly. I don't want to wait for a power failure to find out whether my services still recover from a cold boot.
- toast0 5y agoIf your push process is reasonable, you don't need to confirm boot time config changes on all machines. When you make a boot time config change (which, in my mind, includes kernel version and OS version to some extent), decide if it's nice to have or needs to have. If it's needs to have, well everything gets rebooted and no badges for a while. If it's nice to have, group hosts by kernel version and hardware spec and reboot one or two hosts with median uptime; your config will be tested, and you can still earn badges. If you can't tell what configs are boot time or not, just reboot a median uptime host in each group once a week. If you can't tell what's important enough to reboot everything, I guess you can just reboot everything every 49.7 days, but you'll never get badges that way. You also won't get production data comparing current kernel to older kernel to see if any new software would have worked better on older kernels, so you have some data to start with when you hit new bottlenecks.