3 ms·
The short version: the file system got corrupted. The backup was just a file-sync over a firewire network to another machine. Meaning the bad data was backed up
by lbrandy 18y ago
The short version: the file system got corrupted. The backup was just a file-sync over a firewire network to another machine. Meaning the bad data was backed up and presumably overwriting the older, good data. They had a RAID but the problem was a software filesystem so the errors just got stored.
He seems to understand how terrible of a design decision he made in regards to the back-up system, and he appears physically affected when having to admit, publicly, the details of the infrastructure (or lack thereof) that caused this.
- moe 18y agoA one-liner to add insult to injury: sed 's/rsync/rdiff-backup/g' <bin/my-backup.sh >bin/my-real-backup.sh
- moe 18y agoAfter watching the vid I have to take that back. Apparently there was a SQL database involved... But rdiff is highly recommended nonetheless.
- hachiya 18y agoAnyone know how rdiff-backup compares to duplicity? I know duplicity has an option to turn off encryption, if one wants to remove that overhead...
- jrockway 18y agordiff-backup keeps a live version of the filesystem available, in addition to backups. This means a full restore is just a `cp -a` operation. FWIW, I've used both, and I like the opacity of the duplicity backups, since I store them on S3. If you are syncing to a nearby disk, though, then you might like rdiff-backup better.
- moe 18y agoSeconded, both have their place. I have tried pretty much all of them, incl. snapBack, dirvish and various homegrown scripts building on top of rsync, rdup, rcs and so on. rdiff and duplicity are the most mature of the pack which shows mostly in their handling of corner cases (connection loss during backup, resume of a partial/failed backup, disk full during backup, handling of really large trees) but also in overall convenience and robustness (legibility of on-disk format, configuration sanity, tools to find a specific revision of a file, flexibility in retention/purge intervals etc.). I generally recommend rdiff as the default tool for backups to a remote spinning disk. duplicity, as parent suggests, is good when you need your archive to be a single large file which helps with handling in some situations. There is also dar worth mentioning which is less useful for incremental stuff but can add redundancy to archives which is good for archiving to unreliable/decaying media (DVD, Tape). Be aware though that older versions had problems with large archives, use a recent version. And last no least, if you have a tape library then Bacula is a mighty tool. Easier to use and pretty much on par in terms of features compared to the commercial offerings and the residents like Amanda. We generally deploy a single backup server here with lots of disks that pulls snapshots from everywhere via rdiff and either mirrors the local repository to a remote location or feeds the precious data to tape via Bacula.
- Maro 18y agoI think there is an implicit assumption here that if you copy over a live Mysql DB file (I assume that's what he was doing), you get a consistent (aCid) view of your DB. This is a false assumption. If this is what he was doing, then possibly the older 'good' backups weren't any good either.
- lennysan 18y agoA great lesson for every saas service, and how a little bit of transparency goes a long way: http://www.transparentuptime.com/2009/02/magnolia-downtime-saas-cloud-trust.html http://www.transparentuptime.com/2009/02/magnolia-downtime-s...