4 ms·
5-nines uptime means anticipating that the components (hardware and software) you bought from another company may fail in devious, correlated ways. Failure to
by ericseppanen 10y ago
5-nines uptime means anticipating that the components (hardware and software) you bought from another company may fail in devious, correlated ways.
Failure to do so is simply incompetence.
Maybe the customer didn't read the fine print of their service guarantees, and that's on them. I would hope that doesn't happen often-- it would be very silly if service guarantees fell apart every time some piece of shoddy equipment (purchased and operated by the service provider) turns out to be at fault.
- txutxu 10y agoThis. The higher SLA system that I've created, was for a military project. Physical network layout: I did choose a double port star topology, this is, every HP5300 modular swith, was connected to each other switch, with two "teamed" ports. Usually if you only do the connections, you got a network loop. But with STP and VRRP in the HP 5300, I got an "always on" network. The network did expand to +50 wifi access points with another proprietary wifi controller which also did have the capacity to gracefully failover connections from the APs, on network splits. Servers behind the switches, where replicated in each segment. So you could at any time turn off (in order, or cutting the power switch suddenly) any rack (switch + servers) and the system did continue to work flawlessly. This was 5 identical racks/switches in star topology. Acceptance tests did include to literally cut cables, literally turn off the UPS and RACK power, etc. We could upgrade any firmware (the HP5300 cabinet, their hot swapable modules, or the servers BIOS/network-card/hard-disk firmware), without any service loss. I'm happy with that result. I've to say, all my other projects where I've work (+15 years), didn't have resources, neither did give any importance, to the needs of network firmware upgrades or downtime (because of a bomb?). In most cases, it was not because technical issues or handicaps, it was because of management. Some projects did listen to me, and did contemplate the issue and planned it as a "maintenance window", or as what today is called "immutable infrastructure": prepare the new one, stop the service, replace, bring up the service. I never did upgrade a switch/router firmware at $job, without having a backup switch ready and pre-configured, in case something went wrong. And preserve the backup one for a prudential time. In my "always on" military project, firmware did need to pass acceptance tests in environments equal to production, before go to any production environment. Edit: remove duplicated info
- godzillabrennus 10y agoI used to do implementation of infrastructure on a project basis and found that every customer wants the best uptime possible until they get hit with the cost. Very quickly active/active failover proposals would turn to active/passive proposals with expectations of downtime if a failure occurred. This web company promising such little downtime was stupid. And now they are bankrupt. Good.
- ChuckMcM 10y agoI remember explaining to a storage customer why they needed three very expensive switches at the core of their network. One to handle the traffic, another to handle the traffic when the first was busy, and a third which was in the rack, patched and updated to one release behind the main switches, that could be swapped in by moving cables in under 5 minutes of operator time. They initially thought I was kidding but then we did the fail tree on a white board. It was an interesting experience for them, understanding what it took to get what they took for granted.
- jerf 10y agoYou used the term "fail tree" which sounds interesting and I tried to learn more on the Internet but hit problems. Do you mean fault tree? https://en.wikipedia.org/wiki/Fault_tree_analysis https://en.wikipedia.org/wiki/Fault_tree_analysis If not, do you have a better search term and/or link I could follow up on?
- ChuckMcM 10y agoFault trees seem to be reasonably close. While developing my thinking and understanding on reliability analysis I came at it from an analog of decision trees. There were a couple of influential talks that got me started, one was by Sandia labs discussing the ways in which they insure that nuclear devices can't be detonated without approval (big requirement) through a process called vulnerability analysis, and another on how Citibank worked their network to insure uptime. The Chaos Monkey series from Netflix was quite fun as well. Google did something similar internally with its DiRT exercises (Diaster Recovery). My "fail tree" (and I'll put it in quotes as unique to my conception of them) analysis consists of identifying a failure, the system response to the failure, a time to fix for the failed system, and a guess at the uncertainty on the fix. So for example "switch hardware failure" is a failure, with a fix time that varies based on "replacement part on hand" to "order/ship/install (replace) the entire switch". The first order is failure/fix tree with callouts of down time. The second order is mitigation/cost with mitigation strategies and their cost resulting in a new call out of potential downtime, and the third order is mitigation accelerators and their cost (which shorten recovery to non-degraded mode) which affect cost and possible down time. Much of that you can do on paper, but sometimes you will have to run experiments to see how long things take to fix.