3 ms·
I don’t see why you think “12 hours affected in the last 30 days (98.31% uptime)” Is trying to spin anything. It’s making easier to see the impact over the las
by Angostura 15d ago
I don’t see why you think “12 hours affected in the last 30 days (98.31% uptime)”
Is trying to spin anything. It’s making easier to see the impact over the last 30 days. I agree with the article
- bartread 15d agoOne thing I would like to see is how many of those hours are during normal business hours in my country. Not all hours are created equal when it comes to downtime and my intuition is that most of these 12 landed within my working hours. In terms of impact that then might mean they were down for 7.5% of the time I needed them, or had business hours uptime of 92.5% which is… both not very good and very disruptive. On the other hand, downtime at 4AM would be much less impactful even if it happened every day and added up to more overall downtime.
- swiftcoder 15d ago> they were down for 7.5% of the time I needed them, or had business hours uptime of 92.5% You can obviously compute this for a particular customer, but being a global service, it's pretty much guaranteed that someone somewhere experienced the worse of those numbers
- names_are_hard 15d agoYou can compute them for the average. In other words, the total customer impact is the number of business-hours of downtime across all customers divided by the total number of business hours of all customers. This is important because it's quite possible that the downtime is biased toward the times they have the most active users.
- Anon1096 15d agoBig systems worth their salt already do this as weighted uptime, considering request successes / total requests rather than wall clock uptime as internal SLOs. But these numbers aren't really ever published because it gives away information about your customer base. https://cloud.google.com/blog/products/gcp/available-or-not-that-is-the-question-cre-life-lessons?hl=en https://cloud.google.com/blog/products/gcp/available-or-not-...
- lanstin 15d ago"Failed customer interactions" - if you have a way to actually see requests before they hit your datacenter, e.g. some async third party client libraries.
- swiftcoder 15d agoOn this topic, as a service operator, it's really nice when you also own your SDKs, and have client-side telemetry about failed requests. Gives you a much clearer picture of end-to-end reliability (at least for the subset of customers who opt-in to telemetry)
- deleted 15d ago[deleted]