4 ms·
What actual metric do you monitor for "datadog outage"? Simply have the deploy tooling make an api request/something else? Our playbook asks the deployer to ta
by temp_praneshp 4y ago
What actual metric do you monitor for "datadog outage"? Simply have the deploy tooling make an api request/something else?
Our playbook asks the deployer to take a look at dashboard X, but automating this would be nicer for some of our CD pipelines
- palijer 4y agoWe don't have a metric for checking and automatically locking deploys. We just integrate their StatusPage notifications into our channels. Datadog doesn't go down often enough for me invest time in automating locking deploys based on it. For anything bigger than a regular code deploys we typically have a runbook ahead of time, and in our template we have a manual check for "make sure datadog is operational" that needs to be checked off on the call. Same with with Github, circleCI, AWS, etc all because we got burned once and and in the postmortem identified that a simple "preflight checklist" would have prevented the issue from lasting so long. It's a good sanity check, reading The Checklist Manifesto influenced me here for these. We work in complex systems, gotta make sure all the stuff is in working order before takeoff.