3 ms·
The lack of additional alerts in the Remediation section is a little bit concerning. Adding an alert for serving stale root zone data is great, but I think a fe
by robhlt 3y ago
The lack of additional alerts in the Remediation section is a little bit concerning. Adding an alert for serving stale root zone data is great, but I think a few more would be very useful too:
- There's a clear uptick in SERVFAIL responses at 7:00 UTC but they don't start their response until an hour later after receiving external reports. This uptick should have automatically triggered an alert. It can't have been within the normal range because they got customer reports about it.
- The resolver failed to load the root zone data on startup and resorted a fallback path. Even if this isn't an error for the resolver it should still be an alert for the static_zone service, because its only client is failing to consume its data.
- The static_zone service should also alert when some percentage of instances fail to parse the root zone data, to get ahead of potential problems before the existing data becomes stale.
- RockRobotRock 3y agonot to be glib but you should consider working for them
- throwawaaarrgh 3y agoIf you can see shit falling apart and you're not even inside the org, it's probably a tire fire, and working there will just be stressful
- RockRobotRock 3y ago"shit falling apart" is a little dramatic. They're a big reputable org that always does writeups when they fuck up. tech people appreciate that. I have my own unrelated issues with Cloudflare as a company
- xcdzvyn 3y agoSure. CloudFlare recruiters, DM me :-)
- jrockway 3y agoSERVFAIL might not be a good enough signal for alerting. I definitely see third-party DNS providers returning SERVFAIL for their own reasons; if it's a popular host (looking at you, Route 53), then you'll proxy those through and end up alerting on AWS's issue instead of your own. They might just want a prober; ask every server for cloudflare.com every minute. If that errors, there are big problems. (I remember Google adding itself as malware many many years ago. Nice to have a continuous check for that sort of thing, which I am sure they do now.)
- physicles 3y ago> They might just want a prober; ask every server for cloudflare.com every minute. If that errors, there are big problems. Yeah this was my first thought too. Why don't they have such a system in place for as many permutations of their public services as they can think of? We're a small company and we've had this for critical stuff for several years.
- XorNot 3y agoIn my experience it's because the monitoring system is usually controlled by another team. So you they don't know what they should be testing, and the developers who know how to do it aren't easily able to set it up as part of deployment. Add issues like network-visibility, and you wind up talking about a cross-team, cross-org effort to stick an HTTP poller and get the traffic back to some server - and so going into production without it winds up being the easier path (because it'll work fine - provided nothing goes wrong).
- lopkeny12ko 3y agoThe alerts you suggested are sufficiently obvious that I'm sure the team has already implemented, or plans to implement, them. The public postmortem report is likely just a small snippet of the more interesting remediation actions.
- AndyMcConachie 3y agoThe real problem here is that things started failing on September 21, but no one noticed until October 4th. Why was there no logging when resolvers started failing to load the local root zone on September 21?