3 ms·
I used to work at PD. When I was there, they followed a lot of (all?) of these guidelines. The app’s ownership was spread across teams, so depending which sect
by kenrose 8y ago
I used to work at PD. When I was there, they followed a lot of (all?) of these guidelines.
The app’s ownership was spread across teams, so depending which section was affected would page certain teams. If it was SEV-2 or higher, that would page an IC and the group of primary on-calls to begin triage. Other SMEs were looped in as necessary.
The anti patterns section is quite authentic. PD had very healthy discussions internally about the topics covered (eg, when it made sense to stop paging everyone, how to make people feel it was OK to drop off the call if they weren’t adding anything).
In terms of the actual “how do they page people if PD is down?”, they had some backup monitoring systems that could SMS / phone on-calls directly. As a piece of software though, PD is pretty resilient, so it was rare to have an outage that affected everything so badly that they had to rely on these secondary systems.
- wbronitsky 8y agoI can verify this as well. I worked on a separate team from Ken, but while we were there, we faced a lot of issues that led the team to decide upon and codify these rules. They are pretty tried and true, if only because they came out of a lot of iteration and testing. Hope you are doing well, Ken!