4 ms·
I run a slightly crazy setup for about 200GB of data: * raid1 on all data. This isn't a backup - it's for end-to-end detection of bit flips in (some of) the st
by pslam 11y ago
I run a slightly crazy setup for about 200GB of data:
* raid1 on all data. This isn't a backup - it's for end-to-end detection of bit flips in (some of) the storage path (and at rest). This needs a better solution (zfs?)
* zbackup of all selected files to external hard disk, overnight. This handles the de-duplication.
* duplicity of the zbackup dataset to S3, immediately following. This is usually a small upload - zbackup diffs are tiny and it doesn't touch files it doesn't need to.
* Rate limited so full backup is about a week, incrementals usually only a few MB.
It seems to work well so far, and I'm prepared to do a fresh full backup every year, swapping disks periodically. This general idea is to keep data in 3 physical places: local, external and remote. Local can be recovered from external. External can be thrown away and recreated. Remote recreated.
Wish there was an all-in-one solution with all the checkboxes checked. It feels silly to have to get there with a bunch of scripts.
- pja 11y agoraid1 on all data. This isn't a backup - it's for end-to-end detection of bit flips in (some of) the storage path (and at rest). This needs a better solution (zfs?) Yeah, btrfs or zfs are the better choice for this than RAID1 in the modern world I think. ZFS anywhere it runs natively, btrfs on Linux. (My experience with btrfs is that it’s fine on a server, but on laptops I’ve ended up with unmountable, unfixable filesystems a number of times. To be fair to btrfs I’ve always been able to recover the data though.)
- riquito 11y agoAnecdotal, but I use btrfs on my laptop since august 2014, had a good number of hard power off not related to btrfs and the filesystem never failed me once. I had some problem with space though using intensively docker: in some situations btrfs thought that the free space was finished and I had to manually rebuild the metadata.
- vidarh 11y agoThe Docker-on-Btrfs problems are/were severe enough that CoreOS now defaults to overlayfs on Ext4. I have a bunch of EC2 servers on CoreOS, and we ended up wiping Docker partitions regularly until that switch...
- pja 11y agoIIRC docker (or any other virtual machine filesystem) in a single image file and hashing filesystems mix really, really badly. Docker could use BTRFS / ZFS snapshots instead of using whole filesystem image files though, so this ought to get better.
- FlyingAvatar 11y agoRAID1 will not get you any protection from bit flips unless you have some specialized software doing that. RAID1 will speed up reading by spreading reads across both disks, but data returned by the array comes from one disk or the other, not both. This means there is no opportunity to detect if a bit flip occurred. Adding a checksumming filesystem that is handling the RAID1 in software would solve this problem. (i.e. If it were hardware RAID as opposed to ZFS's RAID-Z, ZFS could detect a bit-flip on the hardware RAID1, but it can't do anything about it since it is not aware of the two physical disks as a redundant source of data.)
- pslam 11y agoYeah the RAID1 part is weak sauce. However, my specific setup mitigates these flaws: * The intention is to detect at-rest bit flips before they progressively pollute everywhere. I don't mind if it's not instantaneously detected on-access. I perform a nightly full scan - so there's up to 24 hours where an at-rest bit-flip may lie undetected, but it won't progress past that. * I use ECC at all cache hierarchy levels possible, and ECC DDR, with active background scrub. So bit flips here won't occur (with vanishingly small probability). * I use software-RAID1 only. The path between DDR, thru CPU, and to storage controller is unlikely to have a bit flip. From there onwards, the data is essentially written twice, so there is vanishingly low probability of both having a bit-flip, except for systematic failure for that bit pattern in two attempts to two different disks and two different controllers. So there's still some places in the stack where errors can be introduced, but the most common areas of fault are either duplicated or covered by detection mechanisms. My goal is to detect, but not correct, transport and at-rest bit flips. I'll discard everything when an error is detected, and use backups. For my next setup I'll probably switch to a filesystem with a better end-to-end error detection story, such as ZFS, and ditch RAID altogether.
- cmurf 11y agoHow do you detect bit flips with software RAID? A scrub can detect differences, but without checksums it's ambiguous which copy is wrong. And in the case of mdadm, which has no concept of a file system, it won't report the affected file. I'm not even sure you get an LBA, but if you do, then you have to go look that up, accounting for offsets, with the chosen file system to get a file. Conversely, Btrfs scrubs will report a corrupt file path if there's a bit flip in a single copy. If there's another copy available then there's just a kernel note that there was a data csum error and the problem was fixed, no file name path.