4 ms·
Five nines is five minutes and sixteen seconds of downtime per year. That is an absurdly small amount of time, relative to an entire year. Saying it's not hard
by packetized 10y ago
Five nines is five minutes and sixteen seconds of downtime per year. That is an absurdly small amount of time, relative to an entire year. Saying it's not hard demonstrates your ignorance in operating a network.
- greenleafjacob 10y agoPartial outages multiply the outage window by the fraction that it's available. So if you have an outage that affects 1% of traffic then five nines globally permits that 1% outage to last ~8 hours.
- walrus01 10y agoYou'd be surprised - I work for an ISP that sells five nines SLA on most of its circuits, and meets it. We spend money to do it. Having 1+1 redundant core and agg router sets at every POP, geographic diversity of inter-metro fiber, full A and B side power/rectifier systems, etc. With the right BGP and ospf design you can absolutely meet five nines availability for an end user customer perspective. We have some places that are approaching six nines. This article reads like they had 1+0 everything and ran into a nasty iOS bug. Running out of ram is amateur hour as well. When we need to do customer service impacting maintenance that will totally take their segment off the net, the hit can be from 15 seconds to a couple of minutes. And that is in a case where a colo customer is single homed to a single aggregation switch like one of our 48-port 10GbE aristas.
- packetized 10y agoWell aware of the pitfalls; I've worked in neteng for most of my career. But your statement of '15 seconds to a couple of minutes' - there's the rub. Do that more than once and you've blown your SLA. Don't misunderstand me, I'm not defending Peakweb, simply saying that running a five nines network is hard - it typically requires deep experience in diverse problem domains.
- walrus01 10y agoIf you read most any SLA - pre announced maintenance with advance notification is categorized differently than unexpected events, for contractual purposes. JunOS and IOS upgrades need to happen, crossconnects need to get moved, agg switches get replaced, etc. Customers know this. It definitely requires ccie level knowledge and at least 10-15 years experience, plus advanced Linux/BSD server admin skills to really do five nines right. It is indeed expensive and requires enthusiastic cooperation from non technical management responsible for budgets. I worked my way up from a level 1 NOC type position, so I'd like to think that after 20 years I have a good understanding of all the possible OSI layer 1 failure modes (and things you can fuck up in configuration at layers 2 and 3), yet the things some other partner and competitor regional ISPs do continue to surprise me.
- vidarh 10y agoIt sounds like the two of you are in violent agreement. If it wasn't hard, people with your level of experience wouldn't be needed to get it right. Nobody here are saying - as far as I can tell - that they shouldn't have done better. But that'd require them to actually have people with sufficient experience and the budgets and buy-in.
- drostie 10y agoThere's lots of hate for Peak Web here, which I totally understand, but it sounds to me like the problem really stems from Machine Zone shopping around for the cheapest possible contract. Maybe I'm just biased because I know that the game in question was basically just trying to sell desperate men discreet shots of Kate Upton's and Mariah Carey's cleavage. I don't know the details of the countersuit, but that one exists at all is pretty telling. It suggests one of many things: it's possible that the contract did not guarantee any amount of uptime; or maybe Peak Web was not adequately informed of the load that the game was going to take on their servers and therefore they want to argue that the level of downtime was reasonable given that they weren't told to expect those kinds of loads; or maybe the contract had a termination procedure and Machine Zone decided to violate that procedure and just drop the company -- which is not something you can just do. I mean, lawyers will argue anything for cash, but it sounds like Peak Web isn't exactly rolling in the cash they'd need to do this on a whim. I don't know what the chances are that the lawyers in question are working on contingency, but it seems plausible.
- mjevans 10y agoWould you also encourage such a network be built with diversity of parts suppliers? I ask because it seems like you'd want to be able to also resist attackers exploiting a bug in a single vendor solution.
- itchyouch 10y agoDepends on the features that are needed. Cisco seems to have a quite a bit of cisco-only protocols that require Cisco gear. And for our particular use-cases, we seem to need those features. Then again, I'm at a ultra low latency shop that requires port-to-port switching at sub 300 nanos and faster. If one is using standard protocols (bgp, ospf, etc) then mixing and matching doesn't really seem to be a problem.
- packetized 10y agoEven OSPF/IS-IS interop is sketchy at best with some vendors, these days. Ah well.
- walrus01 10y agoThere are a great many ISPs successfully using a mix of juniper, Cisco, arista and other stuff. You will find extreme and foundry 1000baseT switches all over the place still. I would never encourage a monoculture of one model nexus 3000 or anything similar to that.
- pinewurst 10y agoThere's a lot of CCIE types out there who feel they have to maintain a totally Cisco shop regardless of equipment or diversity merits. This is finally beginning to fade, but I've seen it more than a few times.
- vidarh 10y agoPersonally I'd never deploy a single model at the very least, when I am given the budgets for full redundancy, but ideally I'd prefer different suppliers too. I've lost too much sleep fixing problems where someone stupidly thought it was "simpler" to deal with the same everywhere. Until the same problem took down every router at once, or the same manufacturing defect caused their drives to start failing at a high rate at nearly the same time... So it's not just attackers, but being susceptible to the same thing triggering the same bug by accident at the same time, or manufacturing defects across a whole batch or model, or affecting a part that's used all over the place.