3 ms·
You can compute them for the average. In other words, the total customer impact is the number of business-hours of downtime across all customers divided by the
by names_are_hard 17d ago
You can compute them for the average. In other words, the total customer impact is the number of business-hours of downtime across all customers divided by the total number of business hours of all customers.
This is important because it's quite possible that the downtime is biased toward the times they have the most active users.
- Anon1096 17d agoBig systems worth their salt already do this as weighted uptime, considering request successes / total requests rather than wall clock uptime as internal SLOs. But these numbers aren't really ever published because it gives away information about your customer base. https://cloud.google.com/blog/products/gcp/available-or-not-that-is-the-question-cre-life-lessons?hl=en https://cloud.google.com/blog/products/gcp/available-or-not-...
- lanstin 17d ago"Failed customer interactions" - if you have a way to actually see requests before they hit your datacenter, e.g. some async third party client libraries.
- swiftcoder 17d agoOn this topic, as a service operator, it's really nice when you also own your SDKs, and have client-side telemetry about failed requests. Gives you a much clearer picture of end-to-end reliability (at least for the subset of customers who opt-in to telemetry)