5 ms·
There’s some cool ideas in here, but you can also just use a git repo that everyone has cloned.
by andyfleming 5y ago
There’s some cool ideas in here, but you can also just use a git repo that everyone has cloned.
- MaulingMonkey 5y agoPhysical copies still have advantages over git in the event of: - Power failure (if your UPS failed, backup laptop might not be charged up / left at home / ???) - Power surge (if your laptops were plugged in, they could've been fried too) - Disk encrypting ransomware (if you're unsure about the spread of the infection, booting a possibly compromised machine might cause more data loss) - Theft (a casual thief might just grab your laptop, and leave the annoying-to-unplug PC tower alone, but that's not a guarantee - boring secretless paperwork is a lot less likely to be stolen) - People forgetting to clone your git repository
- PetahNZ 5y ago- Power failure where all of the above occur at once? - Power surge big enough to take out all the breakers, UPS, servers, desktops, and laptops at once? - Ransomware that had time to take control of the whole network including backups? - Theft of a single laptop? - People, yea I agree with people can be an issue. And all of the above across multiple offices with WFH staff and on call staff? Of all these catastrophic failures it's just as likely to be a fire, or earthquake, etc, where you should have multiple layers of redundancy in place anyway.
- MaulingMonkey 5y ago> Power failure where all of the above occur at once? Sure, it's easy enough. I've had UPSes that only noticibly failed once the power failed, and those outages outlasted my fully charged laptop battery. > Power surge big enough to take out all the breakers, UPS, servers, desktops, and laptops at once? Sure! One lightning strike to an unprotected network link, for example, can threaten to do just that. https://www.youtube.com/watch?v=Ev0PL892zSE&t=354s https://www.youtube.com/watch?v=Ev0PL892zSE&t=354s (note: he even added some protection and another lightning storm still caused damage... so he resorted to fiber!) Even if it doesn't destroy all of your gear, destroying all the main gear your IT crew has login passwords for can slow things way down as they figure out alternatives. > Ransomware that had time to take control of the whole network including backups? Even if it hasn't taken over the whole network, you might not be sure which nodes have been taken over (ransom messages might not appear until a lot of data has already been encrypted), and may wish to leave more sensitive nodes offline for forensics, or to preserve data that hasn't been backed up yet, or to avoid needing to resort to slower restoration from offline backups. Hopefully you have offsite and cold nodes... good reason to have some bootable USBs with recovery images on 'em ready to go as well. > Theft of a single laptop? "but that's not a guarantee" was meant to point out that larger scale theft beyond the typical "just a single laptop" isn't unheard of. > And all of the above across multiple offices with WFH staff and on call staff? Runbooks can potentially be useful for single-office businesses, and means fewer games of telephone relay even for multi-office businesses. > Of all these catastrophic failures it's just as likely to be a fire, or earthquake, etc, where you should have multiple layers of redundancy in place anyway. And runbooks can point you towards those redundancies, many of which may require manual intervention by design (e.g. restoring from cold offsite backups.) You'll note such things as evacuation plans are often printed and displayed prominently near emergency exits, not left in a git repository ;) And, of course, just because you should have multiple layers of redundancy, doesn't mean you actually have multiple layers of redundancy, even if you think you do.
- linker3000 5y agoTo add to this, someone once suggested that the Incident Manager could just print off a copy of the process when needed. Well, first, can you spare the time during a major incident to go print something, or arrange for someone else to do it? Second, I have seen times where a lack of building connectivity stops centrally-audited printing from working. I daresay a fallback should have been configured, or someone (go find them!) has an override, but valuable time could be wasted sorting this out. Third, if the process being followed on-screen needs revising, you shouldn't start doing that in the document 'on the fly' (where's your document change management now?), so out comes the pen and paper...or you could try firing up your favourite editing package to make notes while simultaneously following the electronic process, updating the ticket and chairing the incident response Teams/whatever meeting..or you could just write quick notes in the right place on the printed procedure so that they have context and you can come back to them later during the incident debrief. I am sure there are pros and cons, and some people will be more comfortable having everything in 'electronic' format, because that's how their IT world has been from day 1, but having been in Technical Support and IT Management for some 30+ years (yeah, I'm an old fart), I know from bitter experience what works best under the vast majority of circumstances.
- andyfleming 5y agoThat’s true. If those are the types of incidents you are preparing for, it would make sense to have printed copies. You could also do both.
- fulafel 5y agoHow do you ensure everyone has it cloned (and up to date)?
- andyfleming 5y agoIf you have any network connectivity or are physically together the repo could be shared. So as long as one person has it cloned you should be fine. It’s the same as making sure people have a printed copy. Depending on the run book, it may not need to be perfectly up to date, but you’re right that it may be challenging for team members to remember to update it. One thing that can help is putting run books in a primary docs repo. Then users are making changes and updating the repository more actively. In turn they would likely get updates to the run book just by writing and updating other docs.
- unilynx 5y agoCrontab a git fetch. You can then pull when needed (but won’t have a process overwriting your docs and scripts at the wrong time)