3 ms·
The real learnings of this incident is how to handle incidents effectively. The author cites ones that I use every time I manage one: * Centralize - 1 tracking
by rdoherty 3y ago
The real learnings of this incident is how to handle incidents effectively. The author cites ones that I use every time I manage one:
* Centralize - 1 tracking doc that describes the issue, timeline, what's been tested, who owns the incident. Have 1 group chat, 1 'team' (virtual or in person). Get an incident commander to drive the group.
* Create a list of hypotheses and work through them one at a time.
* Use data, not assumptions to prove or disprove your hypotheses.
* Gather as much data as you can, but don't let a particular suspicious graph lead you into a rabbit hole. Keep gathering data.
If you don't do the above you are guaranteed to have a mess, have to repeat yourself over and over and waste time.
- a012 3y agoI can’t read the blog but Pagerduty provides a good standards for handling incidents: https://response.pagerduty.com/ https://response.pagerduty.com/