4 ms·
Connect your status page to actual metrics and decide a treshold for downtime. Boom you’re done.
by TeeWEE 3y ago
Connect your status page to actual metrics and decide a treshold for downtime. Boom you’re done.
- lrem 3y agoDoes anyone serious do this? That’s an honest question, from a pretty experienced SRE.
- darkwater 3y agoIn a world of unicorns and rainbows, absolutely. In the real world, it's as you probably already know: it's not that easy in a complex enough system. Quick counter-example for GP: what if the 500 spike is due to a spike in malformed requests from a single (maybe malicious) user?
- laeri 3y agoA malformed request should not lead to a 500, they should be handled and validated.
- jon_adler 3y agoTrue, however it also doesn’t impact other users and doesn’t justify reporting an incident on the status page.
- darkwater 3y agoWell, in the real world it might. It should trigger a bug creation and a fix to the code, but not an incident. Now all of a sudden to decide this you need more complex and/or specific queries in your monitoring system (or a good ML-based alert system), so complexity is already going up.
- laeri 3y agoQuery input validation is nearly a solved problem. If you don't I would argue this is an incident if in this case 500's are returned.
- jabradoodle 3y agoYou need to validate your inputs and return 4xx
- darkwater 3y agoYeah and you also shall not write bugs in your code. Real world has bugs, even trivial ones.
- jabradoodle 3y agoIf your service is returning 5xx, that is the the definition of a server error, of course that is degraded service. Instead we have pointless dashboards that are green an hour after everything is broken. Returning 4xx on a client error isn't hard and is usually handled largely by your framework of choice. Your argument is a strawman
- zbentley 3y ago> Returning 4xx on a client error isn't hard and is usually handled largely by your framework of choice. > Your argument is a strawman That's....super not true. Malformed requests with gibberish (or, more likely, hacker/pentest- generated) headers will cause e.g. Django to return 5xx easily. That's just the example I'm familiar with, but cursory searching indicates reports of similar failures emitted by core framework or standard middleware code for Rails, Next.js, and Spring.
- laeri 3y agoIf input validation is not present in your framework of choice then the framework clearly has problems. If you do not validate your inputs properly I am not sure what you are doing when you have a user facing applications of this size. Validating inputs is the lowest hanging fruit for preventing hacking threats.
- jabradoodle 3y agoUsually handled by the framework, you may have to write some code, I'd expect my saas provider to write code so that I know whether their service is available or not.
- tazjin 3y agohttps://www.buildkitestatus.com/ https://www.buildkitestatus.com/
- sjsdaiuasgdia 3y agoStage 1: Status is manually set. There may be various metrics around what requires an update, and there may be one or more layers of approval needed. Problems: Delayed or missed updates. Customers complain that you're not being honest about outages. Stage 2: Status is automatically set based on the outcome of some monitoring check or functional test. Problems: Any issue with the system that performs the "up or not?" source of truth test can result in a status change regardless of whether an actual problem exists. "Override automatic status updates" becomes one of the first steps performed during incident response, turning this into "status is manually set, but with extra steps". Customers complain that you're not being honest about outages and latency still sucks. Stage 3: Status is automatically set based on a consensus of results from tests run from multiple points scattered across the public internet. Problems: You now have a network of remote nodes to maintain yourself or pay someone else to maintain. The more reliable you want this monitoring to be, the more you need to spend. The cost justification discussions in an enterprise get harder as that cost rises. Meanwhile, many customers continue to say you're not being honest because they can't tell the difference between a local issue and an actual outage. Some customers might notice better alignment between the status page and their experience, but they're content, so they have little motivation to reach out and thank you for the honesty. Eventually, the monitoring service gets axed because we can just manually update the status page after all. Stage 4: Status is manually set. There may be various metrics around what requires an update, and there may be one or more layers of approval needed.