3 ms·
In my experience, correctly identifying the culprit is the most difficult, but during an incident, mitigation is the most important. Most incidents I ran into h
by wbsun 6y ago
In my experience, correctly identifying the culprit is the most difficult, but during an incident, mitigation is the most important. Most incidents I ran into happened after a job rollout, in order to find and mitigate an issue, a distributed system needs monitoring, replicated drainable services, rollout canary/rollback. With these, an SRE oncaller doesn't need as much knowledge of the system as the devs in order to handle an incident.