12 ms·
Million times this. Its shocking how "elevated rate of errors for specific endpoint" in your cloud provider status page is actually amplified to be a soft-outa
by altmind 7y ago
Million times this.
Its shocking how "elevated rate of errors for specific endpoint" in your cloud provider status page is actually amplified to be a soft-outage of your product when your writes to disk never return, your databases returning inconsistent data or your orchestration taking some drastic measures for the failing health check.
When you have a lot of components in your cloud mix, failure of one stage(network->balancing->quering->rendering->persistence) bring everything down.
if 10 of your cloud services each have a reliability of 99.999, all together the reliability is not 99.999.
cloud providers can claim mountain-high availablity whereas users will never get their apps running with advertised reliability for now there is multiple subcomponents that can fail.
- altmind 7y agoThe fact that many status pages are updated manually and any incident disclosure need to get approval from management(aws?) does not add to the status page trust. Uptime and error metrics are technical and should be kept away from managers.
- scottmcf 7y agoThis was extremely evident in the slow response and poor communication during the recent Salesforce outage.
- hinkley 7y agoMaybe it's time for a consumer watchdog group to step in and do their own reporting for services like this. Like https://www.isitdownrightnow.com/ https://www.isitdownrightnow.com/ but with sharper teeth. I'm not sure how much time I have to participate but I wouldn't mind chipping in a bit on a co-op in this space. But it might be easier to convince Is it Down Right Now to grow some fangs, or socialize the idea that it does (perception counts for a lot).
- Annatar 7y agoOr it's time to go back to one's own infrastructure and take one's own destiny into one's own hands, along with the responsibility.
- s_Hogg 7y agoI can't remember ever seeing this work out well lol. Happy to be proven wrong one day, though.
- sombremesa 7y agoYou may have heard of a company called Amazon.
- hinkley 7y agoWhat thread do you think you are replying to?
- oarsinsync 7y ago> own infrastructure and take one's own destiny into one's own hands, along with the responsibility. Amazon did this. It went pretty well. So well they decided to sell the results of the expertise.
- s_Hogg 7y agoIt appears I had forgotten about that back-story!
- Annatar 7y agoI have my own datacenter. Works well and has for the past 20 years. Costs me peanuts because I know what I'm doing and how to do it. Will never be an Amazon or any other "cloud" provider's customer.
- masukomi 7y agohow do you address the central "something went boom" issue of the post. Like if a backhoe takes out the line to connection to the data center do you still have any "nines" ?
- dragonwriter 7y ago> The fact that many status pages are updated manually and any incident disclosure need to get approval from management(aws?) does not add to the status page trust. Automated status page updates can also reduce trust, since then the status page is itself exposed to more kinds of system failures.
- hinkley 7y agoSaucelabs is very bad this way. We have tunnels flake out once in a while (I'm still convinced there's a concurrency bug in their tunnel implementation, based on missing events I've seen in test logs), but sometimes Sauce is just having issues. When I'm seeing 100% failure rate, there's often nothing on their status page. Or there's some bullshit metric like VM acquisition times are double normal for, say, some Windows VM. But I'm not seeing 8% failure rate. I'm not seeing an extra 30 seconds. I'm seeing 100% failure rate, with long timeouts, and retries.
- m463 7y agoI worked at a company once where each bug had a really interesting field: root cause I wish I could remember the values you could fill in, they were very intelligently chosen. What I learned: if you didn't know what the root cause was, you probably didn't fix anything.
- finnthehuman 7y agoWas it like an list of predefined values? Where I work they do root cause analysis for everything, but with freeform answers so what you describe might be different from what I'm used to. In general, I'm so used to RCA and layered mitigations (what one of our greybeards calls "belt and suspenders") that I don't know how quality happens without it. I'm a convert to the idea that if you can't fix a problem directly, the fix has to isolate or be as close to the problem as possible. Otherwise the bad state just ripples outward as complexity.
- m463 7y agoUnfortunately this was probably 20 years ago. I know it was a list of predefined values in a drop-down, but I'm not sure if there was a other/write-in field. The gist was that the causes were appropriate and educational. Folks couldn't choose "user is an idiot", instead having to choose "the interface was confusing".
- baud147258 7y agoWell "the interface was confusing" doesn't really rule out "user is an idiot", but most likely will make matters worse.
- JohnFen 7y agoI like having a set of broad predefined values (it helps with standardization, which helps with searching in the future). But a freeform report is also necessary. How else are you going to adequately explain what, where, why, etc., the root cause was?
- devdas 7y ago
- ken 7y ago> if 10 of your cloud services each have a reliability of 99.999, all together the reliability is not 99.999. It's like an episode of Dirk Niblick: https://www.youtube.com/watch?v=bCoGMYV3UPk https://www.youtube.com/watch?v=bCoGMYV3UPk
- vincentmarle 7y ago> if 10 of your cloud services each have a reliability of 99.999, all together the reliability is not 99.999. (The answer is 99.99)
- dragonwriter 7y agoIt's probably not in practice, since that assumes failures are perfectly independent. If they are perfectly correlated, the answer is still 99.999. For most real cases, it will be between those extremes.
- titzer 7y agoI think it depends on how you define availability. 1. Suppose we define availability as "at least one is up". If the failures are completely independent, then the probability of any one being down is 10^-5 (five nines) and the probability of all 10 being down at the same time is (10^-5)^10 = 10^-50 (fifty nines). 2. If we instead define availability as "all 10 are up" (which is essentially equivalent to one failure causes a cascading failure) then in the same scenario where failures are independent, this is (1-10^-5)^10 ~= 99.99% (four nines).
- spockz 7y agoAlthough I agree with your assessment that the total uptime of the system is the multiplication of the individual systems, that doesn’t appear to be the point of the article. > The problem is that they weren't monitoring from the customer's perspective. Had they done that, it would have been clear that oodles of requests from some subset of customers were failing. They would have also realized that certain customers had all of their requests failing. This is saying that if you are small, all your failing requests are within the 0.001% that the provider is allowed to fail. I suppose this depends on what how 99.999 uptime is defined in the SLA.