4 ms·
Water flows down the easiest path. Make runbooks (and their maintenance) easy. This requires management buy-in to create incentives that make runbooks the easie
by dfinninger 4y ago
Water flows down the easiest path. Make runbooks (and their maintenance) easy. This requires management buy-in to create incentives that make runbooks the easiest path.
In my experience, getting management on-board is the most difficult part since engineers will be lazy (I mean this here in a positive, efficient way) and do the easier thing.
I’ve set up runbooks on a few different teams at different companies and I’ve found some strategies that help. Feel free to pick and choose what would make sense from this list:
- Give runbooks (and on-call) time. Subtract the number of expected on-call weeks from your quarterly plan. This is where management buy-in is the most important since on the surface it impacts “throughput” or whatever metric is in vogue. As an aside, it’s unfortunate that most managers (IME) don’t track morale or operational burden as metrics…
- When an engineer is on-call they can do housekeeping, sweep runbooks, dashboards, etc. You want the on-call engineer to have time to do RCAs, and make reasonable changes in response to incidents. Downtime can be spent cleaning and updating. We’ve called out “operational improvements” during standup.
- Every alert links to a runbook. No alerts can be merged without a runbook. Dashboard, etc, are linked from the runbook. The engineer that created the system/alert might not be the first to encounter it at 3am and the person on-call should know what to do.
- Every time you get an alert, scan the runbook. Even if you are a system expert. Something may have changed or someone may have modified a process to be more automated. You could be pleasantly surprised that an alert is now easier to manage due to someone else’s operational improvements that you forgot about.
- If a runbook is wrong or out of date, fix it right then or flag it as a problem and bring it up during standup. The runbook is a part of the software “package” and issues should be treated like bugs that should be addressed.
- Do an on-call review and track the on-call engineer’s sentiment along with metrics around how many wake-ups, weekend alerts, etc. up-to-date, clear runbooks make on-call much less frustrating for systems the on all engineer didn’t directly write. It’s also a good forum for a team member feels they are taking on more of the operational upkeep.
- Call out and celebrate excellent, up to date runbooks! If you were on call and a teammate’s runbook was great and let you effectively manage an incident in a part of the system you don’t have as much experience in, let them know!
Runbook development/maintenance is a part of software development/maintenance (IMO). If someone isn’t pulling their weight with runbooks, then they aren’t doing all of their job. I see it the same as someone who only wants to work on greenfield projects. There’s some percentage of shit you need to deal with as a software engineer, operations is one of them. At least for online systems.