4 ms·
> Bad: HTTP 400 > On the other hand, HTTP 400 level errors mean the client screwed up. This is bad general advice. HTTP 4xx errors mean the client screwed up,
by cle 5y ago
> Bad: HTTP 400
> On the other hand, HTTP 400 level errors mean the client screwed up.
This is bad general advice. HTTP 4xx errors mean the client screwed up, OR you screwed up (a change that e.g. increases 404 rate due to eventual consistency, returns 404 for all content, breaks auth, returns the wrong status code, etc.). Either way the content is inaccessible to the client, the person visiting your website doesn't care if they get a 404 or a 502, they care that the content is inaccessible. Once you get high enough traffic, monitoring 4xx rate is pretty critical to making sure people can actually use your service. (Or monitor the inverse, i.e. a floor on 2xx rate instead of a ceiling on {4,5}xx rate.)
- breischl 5y agoIn the context of alerting, I agree with TFA. You should not be alerting on bad request errors, because you might have no control over it. That said you might want monitoring on it so you can check if the rate jumped at some important point (eg, after a deployment) but I wouldn't look at it on a regular basis. I had something like that on an internal system. The 400 rate would jump all over the place because our edge systems had shitty input validation, and bots would crawl us with broken requests ("can I reserve this item starting last week?" kind of thing) with no rate throttling. After a few years the edge validation (and bot detection) got better, but alerting on that would've been worse than useless.
- cle 5y agoYeah I agree that false positives are a risk with monitoring 4xx rate. I've never seen a satisfactory "bulletproof" way to monitor 4xx rate, it's inherently difficult and simultaneously important to monitor. It's easy in retrospect to say "oh that was a waste of time because it was just bots" but you don't know that until you investigate. I ask myself "if I see elevated 4xx's, at what point do I start to care if they're caused by a bug?" and set monitor thresholds somewhere around there.
- 8note 5y agoThe client who's calling wrong might want some help getting it right. Its empathetic to keep an eye on surges in 4xxs
- closeparen 5y agoSure, but there’s a big world between keeping an eye on surges and waking someone up if it goes out of 3-4 nines.
- raynorelyp 5y agoAuthor here. This.
- k__ 5y agoThis. In my experience, 4xx and 5xx are only valuable to find the right place to look, but in no way do indicate if client or server failed.
- advisedwang 5y agoMonitor the general 5xx error rate so you have high SNR. Cover mistakes with robust probers that should get 200 and then alert on any non-200 response.
- mrmincent 5y agoAbsolutely, I used to work for an online bookmaker and if we saw a spike in 404s it usually meant one of our sports traders had stuffed something up and took down a market early, or a release went out that broke our navigation. In a business that is inherently spikey (i.e. the majority of bets came through _just_ before an event started) we had to be pretty careful about what spikes were good and what were bad.