6 ms·
Creating backup storage sucks
- akersten 1mo agoWhat the author describes here is hard because it's "simple." What's simple because it's "hard" is replacing parts 2 & 3 with a network appliance like TrueNAS running a zfs pool that syncs to backblaze every night. Yeah you have to learn a bit but it won't fall in weird ways like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice. My 2¢
- smarmelling 1mo agoI agree I think one of the main things I learned from all this was that I should probably buy / set up a real NAS. I’ll look into zfs pool thanks for the comment!
- xbar 1mo agoI liked the post because it tells a true story about one of the remaining problems that are hard to solve well without a 3rd party.
- XorNot 1mo agoFair warning with ZFS: do not turn on deduplication at the moment. There is a straight up data loss bug in 2.4.3 (zeroed out files, totally silent). https://github.com/openzfs/zfs/issues/18366 https://github.com/openzfs/zfs/issues/18366 This caused all sorts of grief at day job where dedupe was useful for a big cache. Conversely I've run ZFS at home for like 15 years at this point without trouble. But this one is an absolute nightmare.
- fc417fc802 1mo ago> like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice I think the author is having trouble because he is conflating concepts and roles that should be distinct. Sync, rotating snapshots, and deduplicated backups need to be kept entirely separate if you want any hope of maintaining your sanity. So he's got sync but he's missing some sort of rotating snapshot system which would solve the stated concern of guarding against syncthing replicating corrupted data. Such automated snapshots can then be used as the source to feed the backup pipeline. That hard drive doesn't make a good backup because it seems that it is always online. You need an offline backup that you manually plug in to run the job once every so often. He's also making this more difficult than it needs to be by insisting that the backup drive be compatible with windows. Plug the drives for both snapshots and backups into a linux box, format them with a modern filesystem, and get on with life. Sync is its own clusterfuck and I have yet to arrive at a satisfactory solution myself despite wasting inordinate amounts of time on it. IMO you either go with a network share or you make due with the "least bad" option of syncthing. Personally I've more or less settled on sshfs at this point not because it's particularly good but because it works well enough and doesn't add any additional complexity. Personally I use btrfs snapshots on all my devices, those get streamed across the network to a NAS, and the contents of the NAS are periodically (every few months) stuffed into borgbackup on redundant offline devices. Aside from sync the other problem you'll run into if you're a data hoarder is how to split backups across multiple drives once you exceed a few TB. Because external drives only get so large but the NAS will inevitably keep ballooning.
- nolist_policy 1mo ago> Sync, rotating snapshots, and deduplicated backups need to be kept entirely separate if you want any hope of maintaining your sanity. Not if you use git-annex.
- fc417fc802 1mo agoHuh, I'd heard the name before but I hadn't realized how capable it was. Unfortunately when it comes to a data hoarder such as myself: https://git-annex.branchable.com/scalability/ https://git-annex.branchable.com/scalability/ > Scaling to hundreds of thousands of files is not a problem, scaling beyond that and git will start to get slow. So that's probably insufficient for me by at least a couple orders of magnitude. I'm able to maintain my sanity because snapshots simply capture device state, the NAS collects all snapshots while maintaining their independence, and (so far) borg has been sufficiently scalable to deduplicate any collection I've thrown at it.
- cm2187 1mo agoLess hard these days. AI is a game changer for learning new technologies. It's like having a highly paid expert available to answer all your questions about your little USB backup. Makes learning how to use properly a new software trivial. And priceless when troubleshooting.
- drdexebtjl 1mo agoIt sounds like most of your frustration can be eliminated by actually using your NAS as the ground-truth for all data, using something like SMB/NFS instead of SyncThing. Then you only need to backup the NAS.
- smarmelling 1mo agoOh interesting I hadn’t considered this. I guess the only trouble would be if I am out with my laptop and I have no internet connection but this seems like a good tradeoff for simplicity
- Hamuko 1mo agoApple's Time Machine still backs up your Mac even when you don't have access to the backup target (USB isn't connected, network isn't available). It just stores the backup information on your local storage and then transfers it over to the actual backup target when it's back online. Don't know if there's similar solutions for non-macOS systems though. https://support.apple.com/en-us/102154 https://support.apple.com/en-us/102154
- drdexebtjl 1mo agoRestic can backup to a local repository and later copy the snapshot to a remote repository. But if you’re on Linux, it’s better to use ZFS or Btrfs and native filesystem snapshots, since they’re atomic.
- zenoprax 1mo agoBeware of interactions with git and syncthing. It's fine for straightforward repos where all you do is commit but as soon as you start doing more complicated branching and re-basing you will start to generate lots of `sync-conflict` files. I haven't really found a reliable way around this so I've decided to just manually rsync from my desktop onto my laptop when I want to work remotely (or more recently, SSH into my desktop directly instead and work off that).
- 1mo ago
- aidenn0 1mo agoWhy wouldn't you use NTFS for backing up a windows system?
- tehbeard 1mo agoIt's a pain in the ass on new systems if it copied restrictive permissions to reset/gain access. Ntfs equivalent of Chmod/chown is a "go have a long lunch" type of operation.
- fc417fc802 1mo ago> One of the other materials that I could not figure out how to back up properly is emails. The reason is that there is no clear way to back it up systematically Doesn't that depend entirely on the storage format? For example if you use thunderbird as a client it has database files that you can copy. Since you say that one of your devices is running arch you should check out some of the email options in the repos. There are at least a few of options that can act as a middleman to bridge between your external accounts and your clients. > So having a plan to roll back to a stable state might be helfpul. I installed and set up Timeshift for this Snapshots make this trivial. Particularly in the case of btrfs if you configure your bootloader to always use the default subvolume then rollback becomes just `btrfs subvolme set-default`.
- ifloop 1mo agoor use imapsync to pull emails, then backup the resulting folder.
- sdcfgy 1mo agoimap-backup. It just extracts everything from imap server as mbox. Then that's rdiff-backup'ed to target disks. You can imap-backup it back to another IMAP server if you need to. However it's better to just delete most stuff so you don't have to back it up. Think my mailbox is about 50Mb.
- bblb 1mo agoOne thing I would add to a modern backup strategy: a deferred offline copy Sure it's bothersome to use the sneakernet once a month (manually copy data to USB stick and stash it somewhere), but when that ransomware hits (can happen on personal devices also), it literally can't access the offline copy, and try to encrypt that one also. Immutable backup solutions are code and thus eventually defeatable from the remote attackers point of view. Offline requires physical access, and maybe a big enough wrench to get the USB stick's disk encryption PIN out of you.
- bob1029 1mo agoTape + Iron Mountain is difficult to beat for offline copies. Surrendering just a little bit of absolute control over the physical media can dramatically increase the likelihood that it survives whatever apocalyptic event. I don't know that I've ever heard of tapes being stolen from Iron Mountain. Even if this happened, what are the odds that they would get all the tapes. Your tapes? It probably looks like the warehouse from Raiders of the Lost Ark in there.
- BLKNSLVR 1mo agoI have a locker at my place of work where I store a few HDDs and USBs. These are the most up to date, but I also have others at my parents place and my in-laws. Gets troublesome keeping track of which ones are up to date as of what date. Good challenge for staying organised though, I've got a whole naming system and numbered hierarchy and scripts that run ordered by priority. Well overdue for a refresh.
- vladvasiliu 1mo agoI do this too, but with ZFS. So, since the snapshots have the creation timestamp in their name, it's obvious which drive has the latest data. But I also tend to remember if I went to the office or to my parents' house last.
- NewsaHackO 1mo agoI think there is a system to do this with git annex that will just keep the hash link in a repo but have the file in a offline drive, so you can easily keep track of files that are in cold storage. It may need you to have your own encryption system though.
- TranquilMarmot 1mo agoI have a server in a closet with a few hard drives attached to it. All of my devices connect to it via Tailscale. I use the wonderful https://github.com/garethgeorge/backrest https://github.com/garethgeorge/backrest as a web UI around restic. Every night, I back the data up to a Hetzner storage box https://www.hetzner.com/storage/storage-box/ https://www.hetzner.com/storage/storage-box/ which is only ~$3.50/mo USD for 1TB of data. I have three "tiers" of data for myself: 1. Important things I care about: these go into a folder that gets backed up to Hetzner (examples include photos in Immich, code pushed to Forgejo, etc) 2. Smaller files I sort-of care about: these usually just go into Proton Drive (examples include design documents) 3. Large files I don't care about losing: these go into a folder on a hard drive that is not backed up (examples include media) This took me about a day to set up and I just check in on it every once in a while. Once a year I will do a practice disaster recovery run.
- martin- 1mo agoWhat does your practice recovery run look like? I've got backup set up, but I don't know how to best test it. For example, do you restore everything from Hetzner, or some random sample?
- microtonal 1mo agoI don't like these staged backup for these reasons. Every step (e.g. NAS backup to Hetzner) can cause an error, so adding steps increases the probability of error. Personally, I just back up directly from my machines to a local storage server and to a cloud object storage with object lock (so that e.g. ransomware encryption or attempts to remove the data do not work).
- TranquilMarmot 29d agoWhat cloud object storage do you use? I went with Hetzner because I was optimizing for price.
- TranquilMarmot 29d agoRight now, I just restore everything to a new folder and make sure that I can stand up Immich, Forgejo, etc. again and point it there and have it working. DB backups for Immich are handled automatically, but for Forgejo I had to write a custom script that does a dump of the DB every night. I'm not sure I could do a partial Immich/Forgejo restore since it would be missing so many files.
- yread 1mo agoJust get a tape drive. Then you will find out how nice and simple disks are
- eviks 1mo ago> It contains at the home directory the folder Sync which is what gets synced across all devices and what needs to be maintained. That's a core Dropbox-level mistake of putting the backup cart before the main use horse - your files should be sorted according to their primary, not backup use > Syncthing works very well, its only constraint is that it does need the devices to be often on if you want consistent syncing. So it's working very poorly, this is a very inconvenient constraint > phone is a whole other beast. It does not have enough storage to Indeed Overall, that fits the title perfectly, and is a bad way to setup your backups. Special needs need special apps.
- exe34 1mo agoWhat solution do you use that syncs/backs up while the device is turned off? IME stuff?
- mbi 1mo ago> The size of Sync is actually fairly small for me, it's just 12GB, but it's not small enough to fit on my 128GB phone. Synctrain is a great iOS syncthing client that connects to your swarm but doesn't actually sync any files (unless you ask it to). You can still browse the contents of your shared folders and selectively download individual files. https://apps.apple.com/us/app/synctrain/id6553985316 https://apps.apple.com/us/app/synctrain/id6553985316
- sdcfgy 1mo agoSome advice from someone who's done it wrong for years and dealt with dead people who have done it wrong for years. Firstly don't try and run a corporate level backup solution or NAS with terabytes of storage and tape at home. You're going to drop dead one day and everyone will hate you for what you leave them to deal with. They will have no idea how to pick through the remains to get important stuff out. It took me a whole 3 days, as someone who knows how this works, to get into a dead relative's NAS appliance and find any mention of any info to get into his EV which had basically bricked itself in the driveway while he was in hospital. Secondly, delete everything. Not joking. Delete as much as you can. Everything you've watched, delete it. All the old emails, delete them. Tracks you don't like on CDs you've ripped, delete them. Old software ISOs, delete them. Thousands of photos of the same thing taken 1 second apart, delete them. Thousands of PDFs of books you will never read, delete them. I went down from a whopping 8TB to 300Gb doing that and it's still shrinking monthly. Eventually it'll fit on one mediocre computer and that makes life so so so much easier. Then look at your backup strategy. It'll look simple then. Mine is three external disks. 1. 1TB T7 shield (not encrypted) which lives at my partner's house and gets an rsync once a month. 2. 1TB T7 shield (encrypted) which lives in my bag and gets an rdiff-backup once a week. 3. 2TB Lacie HDD (not encrypted) which lives in my fire safe at home and gets an rdiff-backup once a week. If I drop dead, my partner can just plug the thing into her computer and get stuff off it. I'd avoid the cloud if possible as well. One of the worst situations I've seen is someone who confidently pushed their NAS contents to S3 and then when the NAS blew up they had to pay a lot of money in transfer to get it back again. Hundreds of dollars in fact. On top of the price of a new NAS. That might be a last resort option but it should NEVER be the first line backup. Don't make your life any more complicated than it needs to be. I implore you.
- cm2187 1mo agoI don't particularly like the idea of my relatives sneaking into my files when I am dead. I destroyed all my father's files when he passed to respect his privacy.
- someothherguyy 1mo agoeveryone should do this, but it is difficult. imo, it is better to talk with your parents beforehand (if possible) to understand what they want to share with you and what they do not.
- 8fingerlouie 1mo agoI can't tell you what works for "you", but here's my setup, with the hope it inspires someone: - Live data lives in the cloud. I could self host it, but it would always be inferior to the cloud offerings, and usually more expensive for my ~3TB data. - I make local backups nightly - I make remote backups nightly to another cloud. - Once every week I make a local backup to a device that is offline 6/7 days a week (an old Synology NAS that powers on automatically, and powers off when it's been idle for 30 minutes). - Every year I curate our photo library and burn a set of M-Disc Blu-Ray copies of all photos created or modified in the past 12 months. Two identical sets, one stored at home, the other stored at a remote location, each clearly labeled "Photo Backup <year>". - Every year I also update a couple of external HDDs with the entire photo library, again identical copies, contents are verified yearly and updated, and stored again. Disks are also clearly labeled as "Photo backup". As others have written, curate your data. In my case we have a 2.5TB photo library spanning a couple of decades, but that could easily have been 4TB without curation. I only backup documents in the nightly/weekly backups. (Personal) Documents usually only hold value for a short time, and after that it's mostly sentimental. Anything media, or anything downloaded from the internet like books, music, etc, regardless of if I purchased it or pirated it, is not getting backed up. If it came from the internet, there's a good chance it can still be found on the internet. I also (mostly) don't run RAID. RAID is for availability, and since all my important data lives in the cloud, and I will likely survive if one of my backups dies, there's little reason to run raid. The only exception is the share where PhotoSync backs up our photos, which is on a "small" RAID1 volume mainly because it acts as the source of all the other photo backups, so consistency and correctness is important.
- smarmelling 29d agowow thanks for all the feedback - this was very helpful and I hope I'll be able to apply many of this advice in my next iteration of my backup storage set-up.
- justsomehnguy 29d agoNB: Despite explictly mentioning 3-2-1 at the start TFA do not describe a proper 3-2-1 strategy.
- aborsy 27d agoThe best backup solution that I found is ZFS send (even if it’s not strictly a back up tool). ZFS send raw encrypted stream from laptop to a server automatically. Encrypted compressed incremental. No need to verify: if snapshot exists at destination, there is no error. If a problem is encountered, restore from redundancy. Errors can be detected and corrected. Restic has also been working great. The problem is that, it has no redundancy to correct errors. The assumption is that, server holding repository will have redundancy, but then I can replicate directly to that server. I worry that at some point, there will be an error in repository. I may loose 5 years of snapshots (restic has some functionality to remove involved snapshots and rebuild the index, but may not succeed). Kopia has error correction, with Reed Solomon. Borg2 has perpetually remained in beta. Looking forward to test it. I stay away from sync: problems with permissions, git repo, conflict, incomplete sync status in syncthing etc.