4 ms·
That’s bullshit. In reality (I was affected by this and it’s now fixed), this happened 3 days ago, and I kept watching the status to see if they would identify
by nullspace 7y ago
That’s bullshit.
In reality (I was affected by this and it’s now fixed), this happened 3 days ago, and I kept watching the status to see if they would identify it.
I had to upgrade to premium support for them to even respond to the issue. I filed the issue on Friday or Saturday, and they got back to me on Sunday. And it looks like they have fixed it now.
This was not a quick response time, don’t give them credit for this one.
- mk89 7y agoIt's a nice trick that many companies use. The best way is to build small agents to monitor the service you depend on to know whether they truly respect their SLA. In case of LastPass they don't even have an SLA....so good luck with an updated status.
- arcticfox 7y agoI learned another nice trick from GCP the other day; Stackdriver log ingestion was down, at least for me and a number of people on Twitter, and they simply put a yellow warning at the top of status.cloud.google.com while fixing it instead of making an official incident. Magic, 100% uptime!
- ghastmaster 7y agoIf the service is down for a limited amount of individuals I consider it still up. This does beg the question of how many constitutes "down". I think the nature of the problem and quantity of users affected is important.
- deleted 7y ago[deleted]
- jazzkingrt 7y agoI disagree. The SLA ought to be made on a per-customer basis. 1% of users affected would mean 1% of users would be entitled to refunds/remedies that the SLA prescribes.
- antoncohen 7y agoGCP has different SLAs for each product, but the ones I've seen are per-customer. The details vary by service, but they generally define downtime, like x% errors for y% amount of time, and then have financial penalties for z% downtime[1]. There are sometimes clauses requiring retries with exponential backoff[2]. Their SLAs are short and readable. It is worth reading a few of them, especially if you have a SaaS and are thinking about your own SLAs. Just search Google for [google $product sla] to find them. [1] https://cloud.google.com/stackdriver/sla https://cloud.google.com/stackdriver/sla [2] https://cloud.google.com/datastore/sla https://cloud.google.com/datastore/sla
- deleted 7y ago[deleted]
- ghastmaster 7y agoI did not realize there is tracking of uptime for individual agreements. I think the status.cloud.google.com from the commenter I was responding to is a general uptime status for all users, correct? I checked out the GCP SLA and see that it is tracked on a monthly basis which affects billing(as antoncohen points out)
- jrockway 7y agoThis isn't as bad as Slack, where they will acknowledge an incident, but then if you go back and look at their status history a week ago it ends up being understated and they update the uptime to 100%. I know there is always the case where "it's just me", but I'm talking about an incident that was widely reported in the media because it was so widespread. While the incident is ongoing, they do provide status updates... but after a couple weeks pass, the global outage that affects everyone silently disappears from their archive. It's quite interesting.
- nojvek 7y agoYeah slack has some very shady SLA practices. I don’t think we can trust companies to self report uptimes. There kind of needs to be a third party SLA escrow of some kind to really make things work.
- jrockway 7y agoYeah, I've always encountered resistance when wanting to report accurate SLAs. I have found that typically SLAs are based on what the marketing team thinks sounds good, not what the technology can provide, so the goal is to hide outages and write legal documents that say "when we say uptime is guaranteed, what we mean is that we'll give you some insignificant amount of money when you're down for several days." So everything is legally in the clear, but customers assume some sort of reliability that doesn't exist. What this means is that it's a race to the bottom; if one company claims 100% uptime, the next company either has to explain why that's bullshit (and people react negatively to negativity, even when it's true), or do the same thing. The result is that everyone now has 100% uptime. The problem with this model is that it doesn't allow engineering teams to set realistic goals to improve reliability. Because a customer being down for 3 days doesn't cost the company any money, you can't deploy expensive engineering resources to prevent that sort of thing from happening again. Meanwhile, if you track your SLA accurately, and compensate customers in a way that's commiserate with the inconvenience they experienced (something like "the entire month is free if we're down for 8 hours in a row"), then you can start doing real engineering. You have a clear number that shows where you're at now, and you have a goal for where you want to be, and you have a cost associated with that goal... suddenly you can make intelligent decisions about what to work on. This class of outage costs us $600,000 a year. It would take one engineer at $200,000 a year 3 months to fix it. There's $550,000 of free money. Instead of being a cost center, you're a profit center! And customers get a better product. How is that not a win? I'll never understand. One thing I liked about working on Google Fiber back in the day is that US-based telephone support was not something that we would compromise on. It was expensive! So when we could eliminate classes of problems that people call in about, like poor WiFi connectivity or bad TV remote Bluetooth pairing, you could directly see the savings in support cost. You could spend a year debugging WiFi, and instead of looking like flushing money down the toilet, it looked like making money. It was a joy to work on. But obviously a very uncommon way of accounting. It's easier to say "everything is perfect, we dare you to cancel" than to invest in engineering. As an engineer, that's sad; we want the world to work better... but it's only possible with resources.
- Spooky23 7y agoAll providers do this where they can. Office 365 had an issue where their DNS resolutions were fubar and impacted the service for certain customers, but their position was that the service itself was fine.
- hinkley 7y agoSaucelabs are kings of the "100% outage for 5% of our users = 95% availability" status update. In particular I'd see repeatable problems where they couldn't launch whatever browser X operating system in under 2 minutes (when our tests would time out) and list allocation time as 'elevated' (say, 8s average vs their normal of 3s). If you start believing your own statistics you get into almost as much trouble as believing your own PR. For an 18 month period where they were particularly bad, I think they only copped to an actual problem one time out of around a dozen cases where our CD pipeline was blocked for half a day or more unless we just turned off e2e tests entirely.
- OJFord 7y ago> the "100% outage for 5% of our users = 95% availability" status update What's a fair way of producing a single number though?
- Normal_gaussian 7y agoPer user guarantee levels. i.e 100% outage for 1 user is equivalent to a 0% availability guarantee. The relationship between client and service is the same irrespective of the number of clients the service has - a number which means exactly nothing to the client. But more importantly, availability numbers are for informing a client about how much incidents outside their house can affect them, and reasonable courses of action to take when it does. When using the numbers internally, the fudged number is equally misleading. Unfortunately there exist fewer adversarial relationships internal to an organisation to prevent these short sighted statistical nonsenses.
- djannzjkzxn 7y agoI think this is an artifact of how SLAs are tied to billing. Anyone who had ever billed a corporation knows how they will jerk you around. It’s obvious that you won’t get a straight answer if you go and ask a company how much they owe you. That’s what a status page is. It’s the company’s first offer in the negotiation on how much they owe you for the outage. You need to calculate your own number in response. Maybe a more scalable solution would be a third-party company that sells this information. I think there’s a lot of money to be made there.
- mk89 7y agoWell, then what's the point in keeping a status page ? Ah, right. Marketing. That's my conclusion on what status pages have become. Which of course raises the question: what do I do when I see a service with N problems in their status pages over the last X days? Are they being naive, or was their service so bad that they were forced to write it down? I agree with you, there is a lot of money to make there. I think there are already a few companies doing that, though.
- djannzjkzxn 7y agoI think when a service provides its own status page, the customers are less likely to build their own status monitoring, so the service can get away with more downtime.
- agapon 7y agoI also have been having the problem in a browser since last Friday. Fortunately, lpass command line tool kept working without issues, so I never bothered to report. The problem is fixed now. Maybe it was related to the MFA as I have been prompted for it today. FWIW, this was with Firefox and FreeBSD.