4 ms·
> enabling bad cultural practices I strongly disagree. There is nothing culturally bad in a system issuing an error if there is an error. Sometimes systems iss
by gklitz 2y ago
> enabling bad cultural practices
I strongly disagree. There is nothing culturally bad in a system issuing an error if there is an error. Sometimes systems issue errors that are considered noise by supporters because they are not actionable, but forcing a system to not issue an error just because your support team cannot directly take action on it is an extremely odd leakage of team responsibilities and bound to have unintended consequences. Imagine a developer telling management that they didn’t implement error checking on some edge case because the support team told them they didn’t have documentation about how to take action for instance. The appropriate response there would be “why on earth are you asking support permission to add error messages for a known error?”. On the other hand, if a support team is drowning in noisy error messages they need tooling to make it easy to distinguish between those and other messages that need to be reviewed of have action taken.
- remus 2y ago> There is nothing culturally bad in a system issuing an error if there is an error. That's true, but if the error says "PANIC! EVERYTHING IS DOWN" when it's not true, then it's asking for an action that's outsized to the problem. Error messages are fine, but they just need to be classified and responded to correctly, and noisy alerts are typically the ones that are misclassified and demanding attention they (probably) don't deserve.
- madeofpalk 2y agoThe context here is alerts triggering on-call. If the error is not-actionable, why wake someone up in the middle of the night because of it? I don't think anyone is rejecting the observability of these errors, but just that there's no point in having it alert/wake someone up unnecessarily.
- everforward 2y agoThen it either shouldn't be an alert (and instead part of some kind of summary report or some such) or the devs need to take on call. It is an exercise in frustration for everyone to route the page to ops just to make ops call dev; that means dev still has to have an oncall rotation, they might as well just take the page directly. The unintended consequence of forcing alerts down ops' throat is them gradually caring less about pages, because there's a very good chance that each one is unactionable. I've worked places that do this, I've seen it happen first-hand more than once. It starts with frustration and ops being less helpful to devs, and ends in a jaded acceptance where ops people start telling each other "just close it and see if it happens a second or third time, that alert never means anything". At that point, the system may as well not emit the errors anyways because no one is looking at the alerts anyways.