4 ms·
I think you might miss the point here, in the same way as the KDE SAs did. What folks tend to consider the meat of ops work, often boils to a big ole boring ch
by 3amOpsGuy 14y ago
I think you might miss the point here, in the same way as the KDE SAs did.
What folks tend to consider the meat of ops work, often boils to a big ole boring checklist.
The problem is that you shouldn't just elect to skip a whole big section without some seriously good reasoning.
This isn't a slur on you or the KDE guys, hindsight is 20/20. I'm confident though that I'm not alone, that there are plenty of other Ops folks here who read the story and also felt the described setup violated a deep principle and just made feel ill at ease. These failure scenarios are not common, but the do happen often enough that we know to prepare for them.
As an example I'd point to how DBAs handle validation of replication - it's the same principle here.
Just for completeness, an example reason for not having proper restore procedures in place might be 'this is not the prime record copy of the data and it takes less than 24h to regenerate this data therefore this will be out of scope during restore tests'.
- sho_hn 14y ago> I think you might miss the point here, in the same way as the KDE SAs did. Who're clearly not under the impression they can't make mistakes considering TFA is a write-up of design flaws in a mirroring system :). Look at it this way: If you think something obvious was overlooked, then it's good there's another report backing up your point. That's the value in everyone being open about their operations and experiences along the way - you only get better metrics for what works and what doesn't, in practice.
- DougBTX 14y agoYea, it is good to see a writeup like this. I'm sure some of the servers I work with don't have proper backups, but I was cringing all the same waiting for there to be a discussion about why the central git server itself couldn't be restored from a backup.
- Confusion 14y agoI don't understand what you are referring to. What 'whole big section' was skipped? What 'deep principle' was violated?
- DougBTX 14y agoThey didn't have backups for the server, only mirrors of the content (including syncing project deletes to the mirrors). They had to re-build the server and copy in the content from a mirror rather than being able to restore the server wholesale from a backup. Not having a backup is dangerous, you can't recover quickly and risk losing all your data.
- Confusion 14y agoThere is no need for full system backups if you have backups of the relevant data and can rebuild the machine surrounding the data. For instance to restore a buildserver, you run a script that creates and provisions a new VM, clones a git repo that contains the configuration of the CI server, clones the repos it should build and the server is ready to go. No manual actions and no backups needed.
- 3amOpsGuy 14y ago>> if you have backups of the relevant data They didn't have that, or at least not recent enough backups. Synching rather than snapshotting. It's a subtle difference (sync permits delete).
- mpyne 14y agoOn the contrary, the backups were all too recent. The system is designed such that the master repositories should never actually lose objects (even with force pushes and branch deletions, the admins make backup copies of the HEAD branch before letting those run so that the blobs remain in the repo). As it turns out though there are repo tarballs generated periodically which would have served as a perfectly acceptable backup, and some other things the sysadmins could have done. The bigger shock for them was that git clone --mirror wouldn't actually run the git integrity checks (which they had mistakenly assumed).
- 3amOpsGuy 14y ago