5 ms·
I don't mean to rag on reddit specifically, but it seems like another example of traffic light statuses being useless. Github had similar examples. I don't get
by FLUX-YOU 8y ago
I don't mean to rag on reddit specifically, but it seems like another example of traffic light statuses being useless. Github had similar examples. I don't get why anyone uses them.
- rlyshw 8y agoMaybe I don't necessarily understand what you mean by "traffic light statuses" but I feel like the red=down; yellow=issues; green=all-good; indicators are pretty effective at quickly and concisely communicating the status of a service to the general public. I don't see that as useless? Unless you are talking about such a system for the internal infra team. I doubt that they are looking at this public-facing "traffic light status" page, however.
- FLUX-YOU 8y agoTraffic light model means using red/yellow/green. The problem I see with a lot of sites is that there is plainly a problem, but the status is still green. I was checking reddit's status page and Aug 20th was green throughout the entire downtime and still is. The amount of times I've seen it be inaccurate makes me think it is not a good model. I would much rather have a graph of requests served and other metrics (which reddit has). You can very clearly see a dive on the graphs and know something is wrong. This is what I'm referring to: https://i.imgur.com/UAtFnfI.jpg https://i.imgur.com/UAtFnfI.jpg
- regecks 8y agoInaccurate status pages are fatiguing. Why are they not automated? It's such a drag to have to use Twitter and IRC to bully their CSRs into actually acknowledging problems. The cynic in me says that these companies want to get away with not making any public statement about having an outage unless they absolutely have to. That's certainly the way it was when I worked in web hosting, the bosses would take every opportunity to avoid reporting incidents if nobody complained. We've all seen the infamous AWS "all agreen" status page while an entire region is melting. Stripe have been slow on acknowledging their API being broken and then later deleted their eventual Tweets admitting a problem. I've sat in the Linode IRC channel with staff members acknowledging outages for the better part of an hour without Twitter or the status sites updated. What's going on? Your engineers shouldn't have the status site and remediation competing for their time and attention. Either hook the status site up to Prometheus/Pingdom and deal with the occasional false positive, or sort your customer communication out. Infuriating as a user and customer. Makes me really appreciate companies like OVH which have comprehensive public smokeping metrics in every region from many ASes.
- forgottenpass 8y agoThere's this short paper called "How Complex Systems Fail" [1] and one of the points is that "Complex systems run in degraded mode." Basically it means that everything is always partially on-fire at any give time. But keeping a "everything's OK" flag flying is a big part of a marketing image. I hate this too, but I get why they do it. A status page that shows things are constantly 1-10% broken doesn't inspire confidence. For example the AWS "all green" event you mention, misleading public-facing information is the bread and butter of the seemingly never-ending line of people that want to browbeat me about switching to $cloud_service_x. I'm not saying AWS isn't great, nor making a point specifically about AWS, but how do you convince people that they're a better cost:value ratio than having a good ops team that keeps our internal-facing services redundant across the basements of the three buildings closest to me? By pretending like the fires don't happen and hiding the ops cost behind ridiculous hourly hardware rates. [1] http://web.mit.edu/2.75/resources/random/How%20Complex%20Systems%20Fail.pdf http://web.mit.edu/2.75/resources/random/How%20Complex%20Sys...