6 ms·
From my perspective as a business, I don't see a difference between my server being down and all of IRC being down. Either way my comms are down. All I care ab
by andrepew 5y ago
From my perspective as a business, I don't see a difference between my server being down and all of IRC being down. Either way my comms are down.
All I care about as a business is about frequency and duration of those outages. If Slack outdoes whatever my in house solution is, Slack wins in reliability.
You can make arguments about the net impact on society of everything being down, but from my perspective as an individual business, having everything down is a win was well.
If I explain to client that we have a delay because our in house service is down -- fault falls on our business. If I explain to a client it is because Slack is down, odds are they'll blame Slack, not us. This is especially true if they're impacted by the Slack outage as well or the outage makes news outlets.
Outages from services likes Slack, AWS, etc. that are wide spread are almost treated like acts of god. Clients seem to be far more forgiving for those sorts of issues in my experience.
- _jal 5y ago> If I explain to client that we have a delay because our in house service is down -- fault falls on our business. This is an under-appreciated competitive advantage "cloud" services offer - outsourcing blame. Somewhat similar to using staffing agencies as lawsuit shields.
- throwawayboise 5y agoNo I don't think this argument holds water. Whose decision was it to use the centralized cloud service? Ultimately it's on you if your business can't respond to the customer.
- andrepew 5y agoIf the decision appears reasonable in the context of your industry, you won't be blamed. Phrases like "Nobody ever got fired for buying Cisco/IBM/etc." are tossed around fairly often. If you chose the expensive, gold standard for your use-case and got burned anyway, odds are it won't be held against you. Of course there are exceptions to that. If AWS going down results in loss of life -- you should take measures to avoid it. If the stakes are a day of lost work, the cure might be worse than the disease from a cost perspective.
- Spivak 5y agoIt's not about whether it makes perfect sense, it's about how understating your users are. People treat large cloud service outages like storms. When AWS has issues people aren't mad at me any more than they would be if they couldn't reach my brick-and-mortar store because of a flood.
- jcims 5y ago> If I explain to a client it is because Slack is down, odds are they'll blame Slack, not us. This is especially true if they're impacted by the Slack outage as well or the outage makes news outlets. My perspective is skewed b/c I work at a place that invests 8-9 figures in their supplier risk management program, but in that domain you don't get any breaks here. This dependency is discovered during pre-contract due diligence and your plan and test history to continue operations in the event of a vendor outage is assessed in the overall risk profile of your company. It's obvious this type of consideration in risk management is picking up steam and I'm expecting an industry to start forming around it. I'm also waiting for 'chaos engineering' features to start creeping into SaaS products.
- boardwaalk 5y agoWill Slack actually outdo your in house solution though? If you have a service (Slack alternative or otherwise) locally hosted, you're cutting out a whole bunch of failure modes: 1. Upgrades/maintenance at inopportune times (you can control this, or not even upgrade at all if you don't need it) 2. Networking issues, if you can colocate with with upstream or downstream dependencies 3. Many scaling issues (disk fills up, etc) since you're only serving your own needs and you're not trying to put as many users in as little capacity as possible 4. Relatedly, abusive use by other users won't effect you. 5. You can keep the service off the public internet and avoid things like DDOS attacks It's nice to do able to blame someone else, I get, but that seems kind of gross, and actually being able to provide the best service overall is what I'd go for.
- Hamuko 5y agoAnd how many failure modes are you introducing? If your service is off the public Internet, you now have a failure mode where the company VPN going down kills all productivity. And if you're going to have issues with a disk getting filled up, that's most likely going to happen with an on-premise solution than with a hosted solution where they have a team to take care that the disks don't "fill up". And not keeping your software updated because "an update would happen at an inappropriate time and we don't need it" sounds like a good vulnerability vector.
- iso1631 5y ago4 or 5 slack outages in the last 12 months. Happy to do the update at a quiet time (say 3pm on a wednesday before we go to the pub), but wouldn't do it during a major event (say the middle of the world cup final when we're monitoring things very closely)