9 ms·
One problem with file based backups is that they are not atomic across the filesystem. If you ever back up a database (or really any application that expects at
by ratorx 2y ago
One problem with file based backups is that they are not atomic across the filesystem. If you ever back up a database (or really any application that expects atomicity while it’s running), then you might corrupt the database and lose data. This might not seem like a big problem, but can affect e.g. SQLite, which is quite popular as a file format.
Then again, the likelihood that the backup will be inconsistent is fairly low for a desktop, so it’s probably fine.
I think the optimal solution is:
1) file system level atomic snapshot (ZFS, BTRFS etc)
2) Backup the snapshot at a file level (restic, borg etc)
This way you get atomicity as well as a file-based backup which is redundant against filesystem-level corruption.
- pixelmonkey 2y agoI agree with you, of course. On macOS, Arq uses APFS snapshots, and on Windows, it uses VSS. It'd be nice to use something similar on Linux with restic. In my linked post above, I wrote about this: "You might think btrfs and zfs snapshots would let you create a snapshot of your filesystem and then backup that rather than your current live filesystem state. That’s a good idea, but it’s still an open issue on restic for something like this to be built-in (link). There’s a proposal about how you could script it with ZFS in this nice article (link) on the snapshotting problem for backups." The post contains the links with further information. My imperfect personal workaround is to run the restic backup script from a virtual console (TTY) occasionally with my display server / login manager service stopped.
- vladvasiliu 2y agoI run this from a ZFS snapshot. What I want backed up from my home dir lives on the same volume, so I don't have to launch restic multiple times. I have dedicated volumes for what I specifically want excluded from backups and ZFS snapshots (~/tmp, ~/Downloads, ~/.cache, etc). I've been thinking of somehow triggering restic by zrepl whenever it takes a snapshot, but I haven't figured a way of securely grabbing credentials for it to unlock the repository and to upload to s3 without requiring user intervention.
- magicalhippo 2y agoWindows' Volume Shadow Copy Service[1] allows applications like databases to be informed[2] when a snapshot is about to be taken, so they can ensure their files are in a safe state. They also participate in the restore. While Linux is great at many things, backups is one area I find lacking compared to what I'm used to from Windows. There I take frequent incremental whole-disk backups. The backup program uses the Volume Shadow Copy Service to provide a consistent state (as much as possible). Being incremental they don't take much space. If my disk crashes I can be back up and running like (almost) nothing happened in less than an hour. Just swap out the disk and restore. I know, as I've had to do that twice. [1]: https://learn.microsoft.com/en-us/windows/win32/vss/the-vss-model https://learn.microsoft.com/en-us/windows/win32/vss/the-vss-... [2]: https://learn.microsoft.com/en-us/windows/win32/vss/overview-of-pre-backup-tasks#writer-pre-backup-tasks https://learn.microsoft.com/en-us/windows/win32/vss/overview...
- lmz 2y agoLVM snapshots are copy on write and can be used the same way.
- magicalhippo 2y agoAny backup software that utilizes LVM in this way? Ie automatically creates a snapshot and sends the incremental changes since previous snapshot to a backup destination like a NAS or S3 blob storage.
- lmz 2y agoI don't think the diffs are usable that way. They're actually more like an "undo log" in that the snapshot space is taken by "old blocks" when the actual volume is taking writes. It's useful for the same reasons as volume shadow copy: a consistent snapshot of the block device. (Also this can be very bad for write performance as any writes are doubled - to snapshot and to to the real device)
- magicalhippo 2y ago
- _flux 2y agoYou can also use lvm2 and then you get atomic snapshots with any file system (I think it needs to support fsfreeze, I guess all of them do).
- pixelmonkey 2y agoI never knew this. Thanks for sharing!
- Am4TIfIsER0ppos 2y agolvm requires unallocated space in the volume which makes it kind of garbage to use for snapshots
- fulafel 2y agoOnly a little (as much as data will change during the backup). And default filesystems nowadays support resizing downwards so you can make space after initial partitioning.
- Am4TIfIsER0ppos 2y agoYou have to know in advance to not allocate 100% to root and home otherwise you are SOL when you want to make space later. If you're lucky you can disable swap and temporarily use its allocation to do it, providing that is large enough for the changes.
- fulafel 2y agoThis is not the case, as like I said you can shrink either of those filesystems and its container and use the freed space for this. (Also I think lvm doesn't need the volume blocks to be contiguous on the physical volume. So you might have N free space after volume a and M after volume b, and lvm would let you create a new N+M sized volume.)
- Am4TIfIsER0ppos 2y ago
- deleted 2y ago[deleted]
- hashworks 2y agoWhile I do that, is that really the case? I can imagine database snapshots are consistent most of the time, but it can't be guaranteed, right? In the end it's like a server crash, the database suddenly stops.
- lmz 2y agoYour DB is supposed to guarantee consistency even in server crashes. (The Consistency, Durability part of ACID).
- mdavidn 2y agoThat consistency is built on assumptions about the filesystem that may not hold true of a copy made concurrently by a backup tool. e.g. The database might append to write-ahead logs in a different order than the order in which the backup tool reads them.
- grumbelbart2 2y agoThat's why you do a filesystem snapshot before the backup, something supported by all systems. The snapshot is constant to the backup tool, and read order or subsequent writes don't matter. The main difference is that Windows and MacOS have a mechanism that communicates with applications that a snapshot is about to be taken, allowing the applications (such as databases) to build a more "consistent" version of their files. In theory, of course, database files should always be in a logically consistent state (what if power goes out?).