5 ms·
>'Our Windows guys were still working on restoring the master Active Directory pair and DFS cluster TWO DAYS LATER even with a DR strategy in place.' I'd be ov
by incision 12y ago
>'Our Windows guys were still working on restoring the master Active Directory pair and DFS cluster TWO DAYS LATER even with a DR strategy in place.'
I'd be overjoyed if I never had to admin Windows again, but it's a bit unfair to frame that as an OS comparison.
The most basic of direct from TechNet 'best practices' would reduce the impact of a failure of that sort to minutes.
That said, that sort of failure not at all uncommon. Further, I expect such disasters would happen far less often if an existing, poorly implemented Microsoft environment could be remediated into something better as directly as anything unix-like.
- allegory 12y agoI think it's completely fair. Windows is completely non-deterministic due to the insane amount of coupling inside it and the sheer snowball of crap that has stuck to it over the years. In this case, they had tested DR strategy etc in place. Unfortunately when it came to doing it for real (onto the same backup hardware the DR plan was tested against), TSHTF and Windows threw an obscure COM error loading the AD catalog. That required contacting MS 1st line who didn't know what it was so had to escalate it to the AD guys. One obscure registry key change and an ACL change (related to ESE) and it was back. All best practices followed, yet complexity, poor design and secret knowledge crept in and shot the whole process. Stuff like that scares the shit out of me having been on the end of it way too many times now. Many a night have I spent up at 2AM trying to work out why the hell something odd has gone on inside Windows and taken something out in production. Not once in 20 years has a proper Unix (Solaris/FreeBSD) box woken me up or shafted me for hours.
- Rapzid 12y agoYou're just not trying hard enough. There was a Xen hypervisor bug for years that could, when the stars aligned and the moon was full, jump the clock ahead 50 minutes. Nobody knew what was causing it until during a bug session the right guy was looking at the right bit of code and noticed an obscure problem in some in-line assembly. Every system has bugs.
- vezzy-fnord 12y agoThis isn't about bugs, it's about system transparency.
- allegory 12y agoEvery system has bugs, but as the other reply said, quality engineering and most importantly transparency determine the impact. Never been a fan of Xen. It doesn't strike me as quality software. Then again no virtualization solution has to me, yet. I'd rather just deploy all the services to the base machine and use MAC to isolate them.
- shawnreilly 12y agoThe scenario described would indicate that there were delta's between the testing environment (where the DR strategy was tested), and the production environment. It was probably related to OS updates/changes being applied over time, resulting in a configuration that changes. I've always found it good practice to build and maintain a staging environment that mimics the production environment in all aspects. When configuration changes (aka security patch) are needed, they are tested and validated on the staging environment before they are deployed on the production environment. This gives you an opportunity to test and validate the results in a non-production (aka no rules, no SLA's) environment. Part of this involves validating that procedures such as DR will continue to work as expected on the new configuration, before it gets rolled out to production. From my experience, this methodology minimizes scenarios of unexpected behavior in the production environment (aka downtime). I would recommend this methodology (or anything similar) regardless of the OS/distribution you're using.
- allegory 12y agoThat's exactly what was done. The guys we had were shit hot ex Microsoft admins. The DR environment was the staging environment for the patches. Periodically the production kit would be block copied back to the DR environment and sysprepped. Every step to prevent differences between the two clustered environments were taken.