4 ms·
I think I agree with the article, because I just got bitten by exactly this. My API was returning 404 as a fairly common response when a user makes a typo, but
by foreigner 4y ago
I think I agree with the article, because I just got bitten by exactly this. My API was returning 404 as a fairly common response when a user makes a typo, but Twilio's observability code was treating that as a serious error it needed to alert me about.
- eadmund 4y ago> My API was returning 404 as a fairly common response when a user makes a typo, but Twilio's observability code was treating that as a serious error it needed to alert me about. Then Twilio's obesrvability code is broken. Requesting a non-existent resource is not a serious error if resources are requested by hand-constructed URLs.
- shadowgovt 4y agoThis is one of those tricky philosophical points on web admin, because there are two sources of 404s: - random clients on the Internet poking around at stuff, which you have no control over and causes no harm so should not alert - an error in the links or code of the pages you ship to clients, which will cause a consistent degraded user experience and you should be notified about. In practice, the best way to solve this I'm aware of is not via observation of server-emitted 404s at all... It's via a back-channel for clients to report errors (possibly authenticated, so randos on the web can't fudge your stats) and then tie alerting to that back channel. So you don't track 404s to /foo, but you do alert on your clients screaming at you that they got a 404 trying to access /foooo. Of course, this solution requires your end-users to enable JavaScript.
- nrmitchi 4y agoYou can monitor for the second case and react accordingly, you just don't monitor for them in the same fashion as 5xx class errors (which yes, despite what this article is about, I'm still using that term). For 5xx errors you should typically monitor against a baseline of 0. 4xx class errors you want to monitor for substantial deviation from previous averages. This is a good indication if you broke something, vs client behaviour changed. Remember, monitoring isn't always going to "give you the answer", as much as alert you to a possible issue so you can investigate further. Moreso, in a SaaS application, monitoring for deviations in 4xx rate by client is also helpful. An alert on a single client is likely not a system-level issue, but a devaition across all clients likely is.
- wizofaus 4y agoWait - it's exactly the "random clients... poking around" that I'd want the alerts for! 404s occurring due to intentional but broken requests is fine - the caller will know about it and deal with it as necessary. But either way, it is a reasonable argument in favor of not using 404 in the case an endpoint was matched but the specific resource id/path was not. It's not entirely dissimilar to the distinction between "file not found" and "directory not found".
- shadowgovt 4y agoIf you alert on 404s to URLs that nobody should expect to exist, you create a vector by which malicious actors can wake up your oncall staff at 3AM. That's probably not something you want to do to your oncall staff.
- wizofaus 4y agoDepending on what sort of service you're running that might actually be justifiable, but I wasn't assuming a threshold breach for 404s would be waking people up in the middle of the night. At any rate, any such alert is always somewhat vulnerable to that problem, regardless of how the error codes are being triggered. Sounds more hypothetical than likely-in- practice.
- dementiapatien 4y agoDoes that mean Twilio sends an alert every time some random webscraper tries to GET some favicon or /admin path that doesn't exist on the server? Doesn't that happen hundreds of times each day?
- fishtoaster 4y agoThis is probably the biggest reason I like the author's approach: a lot of tools have assumptions about what, eg, a 404 means that might not match what it means as an application error. For example, my API was also returning quite commonly as my frontend checked the existence of various records. As a result, my chrome console was flooded with red error alerts about failing requests (404s), even though each request had "succeeded" just fine. In another case, I had a site that used http basic auth. An xhr api request returned an expected 403, which resulted in the browser suddenly concluding the basic auth (which was unrelated to the api call) was invalid and the user needed to be reprompted for credentials again. Both of these could be argued as browser problems, but that's the point: browsers (and observability tools and many other things) think they know what a given http status means. Using http statuses for app-level errors often breaks that.
- Merad 4y ago> An xhr api request returned an expected 403, which resulted in the browser suddenly concluding the basic auth (which was unrelated to the api call) was invalid and the user needed to be reprompted for credentials again. That sounds like it was a mistake on the part of the api. The browser should only prompt for credentials if the response headers include a WWW-Authenticate header, but that header should only be included with a 401 response (at least according to MDN). https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/WWW-Authenticate https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/WW...