6 ms·
I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine,
by KeitIG 9y ago
I imagine Mr Good Guy at OVH telling some others:
"guys we have a single point of failure in our architecture with SBG, maybe we should...
- naaah it's fine, we do not have time nor resources"
Then shit happens.
edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me (or backup systems, like generators, not being correctly tested... whatever). I am really curious about the future explanation with what happened exactly
edit2: Why all the downvotes? Even the status page of OVH is down, do not tell me it is good design. We are not here to be charitable, but realist.
- matthewmacleod 9y agoThat seems a little uncharitable.
- dx034 9y agoIf the SBG issue really triggered the outage for the whole network I find it hard to believe that no one saw that problem beforehand. They probably thought that this was too unlikely to happen or that there are other failovers but never tested them properly. No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can happen with one product (eg S3 failure) but datacentre switches should work even if the rest is on fire.
- nkkollaw 9y agoWhat is SBG? A CDN of some sort?
- pfg 9y agoIt's the location of one of their data centers. SBG for Strasbourg.
- seszett 9y agoIt's a collection of datacenters in Strasbourg.
- nkkollaw 9y agoAh, gotcha. So, that one datacenter caused all other datacenters to die..?
- seszett 9y agoIt's supposed to be two separate incidents: power going down in Strasbourg, and fiber network equipment going down in Roubaix (the main center of OVH's network) due to a "software bug". It's explained here https://twitter.com/olesovhcom/status/928587258583748609 https://twitter.com/olesovhcom/status/928587258583748609 in French, they might post an English-language translation soon.
- nkkollaw 9y agoThanks!
- zaarn 9y agoIn this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.
- dx034 9y agoSo what? Losing main power is a standard case for any DC. That's why you have generators. Even a generator failure is nothing out of the ordinary. But that no generators in a DC work kind of indicates that they don't test them as often as you would expect. They just announced that they want to be a "hypercloud" provider on the scale of AWS and Google Cloud. I really hope that a power failure in Virginia couldn't bring down all of AWS.
- zaarn 9y agoThe generators did work but they failed. Both of them. I can not imagine they weren't tested. But even the most rigorous testing can never reduce the total failure risk to 0. It seems OVH just got very very unlucky.
- jbb67 9y agoWhen I've seen things like this before, it's often been the switchover hardware that fails, not the generator as such. it's much harder to test that as you don't want to tell your customers "sorry your sever went down, we were just testing if the switchover worked and it didn't"
- vtsingaras 9y agoExactly that, we too suffered a power loss at our DC that was due to faults in the power supervising and switch-over circuits; we test both generators weekly.
- hinkley 9y agoI worked for a company doing mostly on-premises (wireless) carrier software years back but they still wanted our own server room to be 3 nines for reasons lost to me now. They had installed new electrical circuits high on the walls so that we could survive minor flooding incidents. So our lead Ops guy is unplugging half the redundant power supplies and plugging them into the new circuits, but a few critical servers are single PSU still, on UPS units. This is how he discovers that one of our PSUs has rotted and only has about five seconds of reserve power in it. Big outage, no bueno.
- pfg 9y ago> No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. It happened to GCE last year[1], though it only lasted 18 minutes. [1]: https://status.cloud.google.com/incident/compute/16007?post-mortem https://status.cloud.google.com/incident/compute/16007?post-...
- dx034 9y agoBut that's one product, not the whole DC. As I understand the post, other services worked correctly during that time.
- vabene1111 9y agoits OVH: The Hardware is good DDoS protection is good The Prices are high but support/administration does not work well, i have a lot of really weird story's with them, from them plugging in a keyboard in our server to reboot it (without any reason) to taking down a server for a requested maintenance only to notice after 4 hours of downtime that they did not ask their bosses if they were allowed to even perform the maintenance requested (and then not getting permission to do so after another 2 hours ..) For me it feels like there are some really deep issues somewhere in the whole administration that make incidents like this no real surprise Problem is most other providers dont work any better, so ... Everyone makes mistakes, let's just hope they learn from it.
- tyingq 9y ago>ts OVH: The Hardware is good DDoS protection is good The Prices are high The prices are high? Compared to what? Cheap is their raison d'être.
- vabene1111 9y agocompared to other dedicated Server Hardware, not talking about business Cloud Infrastructure, no idea about that. Sry if that caused confusion
- dx034 9y agoWho's cheaper on the dedicated side? Hetzner can be a bit cheaper than soyoustart but wouldn't be aware of anyone else (with reasonable quality).
- zaarn 9y agoI've tried Hetzner but the Network Peering to Telekom is subpar compared to OVH. On Hetz I got about 40Mbps up/down to my local computer while on OVH I can easily load my DSL 100% without issues.
- 9y ago
- switch007 9y ago> - naaah it's fine, we do not have time nor resources" Yup, been there multiple times in smaller hosting companies. It's basically how it goes. They don't get serious about outages until revenue is severely affected and the brand damaged, they don't get serious about security until there's been a big breach or sales are lost because of lack of certification.
- nolok 9y agoCalling OVH a "smaller hosting company" (which you did do indirectly) is rather funny.
- lampington 9y agoUnless they meant "smaller hosting companies [than OVH]"?
- mbrameld 9y agoThe person you're replying to was relating their experiences at hosting companies that are smaller than OVH. How does that imply that OVH is a smaller hosting company?
- abiox 9y agoto me it's ambiguous; it could be either of > in smaller hosting companies [like ovh] > in smaller hosting companies [than ovh]
- switch007 9y agoYes, a comma would have made it slightly clear: > Yup, been there multiple times, in smaller hosting companies. I find it funny that anyone would think I would or could refer to them as small (and get away with it)
- scriptproof 9y agoIt was a power outage. Their own generators did not work. This explain why all is down, but there is surely something bad in their architecture.
- helb 9y agoA power outage in all locations? They claim to have "22 datacenters on 4 continents"… EDIT: The title on HN in misleading, summary from their CEO here – https://twitter.com/olesovhcom/status/928592231807713280 https://twitter.com/olesovhcom/status/928592231807713280
- notyourday 9y agoThis happens all the time. Every single thing that you see in software development happens in network engineering and data center engineering, except that where in development in general senior people who write software are capable of at least guestimating complexities to provide a marginally unified front against the unreasonable expectations of execs, it is pretty much never the case in neteng or dcops as those that develop software cannot write their heads around the complexities of working with physical hardware. It is rather counter-intuitive. In neteng and dcops, "I don't know and I cannot find out. I can only attempt to mitigate what i think might have caused it for next time" is a very reasonable answer to 99% of the "why this happened?" questions because in order to replicate the situation to test the theory one needs to recreate the same problem again on the same scale. This also means that certain things cannot be tested. Most of generator tests are garbage - turning on generator and running it without production load delivered over the transfer switch does not test anything other than that one can turn on a generator and run it. The problem typically happens not because the generator ( also is there the generator or the first and the second generator? Why is there no generator bank for a non monkey-sized company? ) does not start - the problem is because over time transfer switch develops a problem and unlike generators it is not possible to test a transfer switch where in the event of a test failure the customers won't lose power unless the data center is designed from the beginning to deliver A and B powers over separate circuits to every single customer and every single customer has per system ( not per rack ) transfer switches. Of course it costs a lot more money, something that companies are reluctant to spend.
- Gelob 9y agoAny decent sr. network engineer or architect should be able to design you a network and explain the pros and cons, risks, and future scalability. If someone doesn't know why a failure occured and they can't find out then they aren't looking hard enough
- notyourday 9y agoPlease tell me about about this fabulous senior network engineer and where I can obtain a dozen of them.