3 ms·
Standard incident response structure - but you get to play with hours and responsibility a bit more. eg I'm going to a gig on Saturday, can you watch the #alert
by Mandatum 4y ago
Standard incident response structure - but you get to play with hours and responsibility a bit more. eg I'm going to a gig on Saturday, can you watch the #alerts-outage channel?
All of the major SaaS providers have good blog posts and documents on this process. You don't need to automate and build in rostering and shit, just have a Slack channel like #on-call that says "@dave is watching outages, shift finishes at 9AM and hands over to @jess" then at 9AM @jess ack's, and posts the same message with the person who plans to take over next. It's the person who's currently on-call's responsibility to find a replacement if the other person doesn't show up*.
The main issue you'll have is rostering. If you're super early stage, everyone will be happy to do this. Once you have 10-20 employees though, you'll need to distribute this in a fair way so you don't burn anyone out - and make an exceptions process that everyone agrees to (eg @tu worked until 3AM last night on feature for customer X, @sam will take over for them).
* only do this with people who are seniors and are comfortable with having candid conversations, and won't martyr themselves.