3 ms·
Nice to see articles like this describing a company's incident response process and the positive approach to incident culture via gamedays (disclaimer: I'm a co
by quartz 5y ago
Nice to see articles like this describing a company's incident response process and the positive approach to incident culture via gamedays (disclaimer: I'm a cofounder at Kintaba[1], an incident management startup).
Regarding gamedays specifically: I've found that many company leaders don't embrace them because culturally they're not really aligned to the idea that incidents and outages aren't 100% preventable.
It's a mistake to think of the incident management muscle as one you'd like exercised as little as possible when in reality it's something that should be in top form because doing so comes with all kinds of downstream values for the company (a positive culture towards resiliency, openness, team building, honesty about technical risk, etc).
Sadly this can be a difficult mindset to break out of especially if you come from a company mired in "don't tell the exec unless it's so bad they'll find out themselves anyway."
Relatedly, the desire to drop the incident count to zero discourages recordkeeping of "near-miss" incidents, which generally deserve to have the same learning process (postmortem, followup action items, etc) associated with them as the outcomes of major incidents and game days.
Hopefully this outdated attitude continues to die off.
If you're just getting started with incident response or are interested in the space, I highly recommend:
- For basic practices: Google's SRE chapters on incident management [2]
- For the history of why we prepare for incidents and how we learn from them effectively: Sidney Dekker's Field Guide to Understanding Human Error [3]
[1] https://kintaba.com https://kintaba.com
[2] https://sre.google/sre-book/managing-incidents/ https://sre.google/sre-book/managing-incidents/
[3] https://www.amazon.com/Field-Guide-Understanding-Human-Error/dp/1472439058 https://www.amazon.com/Field-Guide-Understanding-Human-Error...
- athenot 5y ago> Relatedly, the desire to drop the incident count to zero discourages recordkeeping of "near-miss" incidents, which generally deserve to have the same learning process (postmortem, followup action items, etc) associated with them as the outcomes of major incidents and game days. Zero recorded incidents is a vanity metric in many orgs, and yes, this looses many fantastic learning opportunities. The end results is that these learning opportunities eventually do happen, but with significant impact associated with them.
- pm90 5y ago> Regarding gamedays specifically: I've found that many company leaders don't embrace them because culturally they're not really aligned to the idea that incidents and outages aren't 100% preventable. So. Much. This. Unless leaders were engineers in the past or have kept abreast of evolution in technology, the default mindset is still "incidents should never happen" rather than "incidents will happen how can we handle them better". This is especially pronounced in politics heavy environments since outages are seen as a professional failure, a way to score brownie points over the team that fails. As a result, you often have a culture that tried to avoid being responsible for outages at any cost, which (ironically) leads to worse overall quality of the system since the root cause is never dealt with.
- Riverheart 5y agoHas does Kintaba compare to competitors like OpsGenie?
- mkopinsky 5y agoCan I ask for your (biased) take on something? My team (5 devs, 10 people total on the product) currently doesn't use any incident response-specific tooling. We have a Confluence SOP for incident response, a page template for RCAs, an #incident-response slack channel, and Zoom but no specific tooling. Just yesterday someone recommended Kintaba/incident.io/OpsGenie/etc, but I don't know if that's overkill for our team. At what point do you think a tool like yours is necessary or worthwhile, as opposed to using generic tools?
- quartz 5y agoObviously biased but I definitely think you can get good value out of a tool like Kintaba at your scale (if you're only using it within engineering it would actually have no cost since we're free for 5 or fewer users!). Kintaba is built to be simple out of the box and allow more depth and complexity as you grow, so initially you might use it the same way you manually use slack today (announce incidents, create a specific incident channel) where your primary initial value is that it makes those motions easier and helps you be more consistent with how you approach incidents and improve recordkeeping, but as you grow you can start to add oncall rotations for your incident roles, automated actions for different incident types, and other things like tagging for better reporting. Feel free to reach out to us at hello@kintaba.com with questions, or even if you'd just like to chat about how to get up and running!