4 ms·
1000 days of uptime is 1000 days of configuration drift since the last time you confirmed your services still boot properly. I don't want to wait for a power f
by addingnumbers 5y ago
1000 days of uptime is 1000 days of configuration drift since the last time you confirmed your services still boot properly.
I don't want to wait for a power failure to find out whether my services still recover from a cold boot.
- toast0 5y agoIf your push process is reasonable, you don't need to confirm boot time config changes on all machines. When you make a boot time config change (which, in my mind, includes kernel version and OS version to some extent), decide if it's nice to have or needs to have. If it's needs to have, well everything gets rebooted and no badges for a while. If it's nice to have, group hosts by kernel version and hardware spec and reboot one or two hosts with median uptime; your config will be tested, and you can still earn badges. If you can't tell what configs are boot time or not, just reboot a median uptime host in each group once a week. If you can't tell what's important enough to reboot everything, I guess you can just reboot everything every 49.7 days, but you'll never get badges that way. You also won't get production data comparing current kernel to older kernel to see if any new software would have worked better on older kernels, so you have some data to start with when you hit new bottlenecks.