3 ms·
It's confusing to me to read this -- you write like you know what you're talking about, but that can't be the case because if you knew what you were talking abo
by davidu 6y ago
It's confusing to me to read this -- you write like you know what you're talking about, but that can't be the case because if you knew what you were talking about you'd understand how the kind of mistake that happened can happen.
As Matthew said on Twitter, this isn't the kind of mistake that happens twice. But for those of us who do operational work, it's easy to see how this happens once. Bit rot in a cab, work orders from someone who didn't know about the legacy patch panel or assumed too much in their instructions and a catastrophe. From now on, photos for every smart hands will probably be a part of the prep-procedure.
- bogomipz 6y ago>"It's confusing to me to read this -- you write like you know what you're talking about, but that can't be the case because if you knew what you were talking about you'd understand how the kind of mistake that happened can happen." That's a pretty odd and flimsy argument that I can't know what I'm talking about if I hold an opinion that differs from your own. I have years of experience with this particular space. Had a MOP been circulated and the maintenance plan properly vetted then the need for visual confirmation because of "legacy patch panel" and "bit rot in cab" would have been identified. Further you never unplug a cable unless you first verify what the ends of the cable are connected to. As I mentioned this is "datacenter operations 101" stuff. This is not some new startup, this a publicly traded company who has been in the game for a decade now.
- davidu 6y agoIt's not the your opinion differs, is that's you are presenting a cognitive dissonance. If you have years of experience then you would understand exactly how this happens, and yet, you think it was avoidable, which is a misunderstanding of operational failures. If this failure was so easy to avoid, it would have been avoided. It was the result of a combination of failures, as all failures in highly-available systems are. It was a failure in bitrot, in planning, in execution, in communication, in reluctance to failover due to a concern about failing back, etc. That said, like Matthew said, this is the kind of failure that happens once.