6 ms·
Hi there! Author here. I created nFreezer initially for my needs, because when doing remote backups (especially on servers on which we never have physical acce
by josephernest 6y ago
Hi there! Author here.
I created nFreezer initially for my needs, because when doing remote backups (especially on servers on which we never have physical access), it's hard to 100% trust the destination server.
Being curious: how do you usually do remote backups of your important files?
--> With usual solutions, even if you use SSH/SFTP and an encrypted partition on destination, there will be a short time during which the data will be unencrypted on the remote server, just before being written to disk and before arriving to the encrypted filesystem layer.
Thus this software nFreezer: the data is never decrypted on the remote server.
How do you work with this?
- rsync 6y agoborg[1] has become the de facto standard for this use-case.[2] It can run over SSH with the borg binary on the remote server or it can run in an SFTP mode with nothing installed on the destination. [1] https://www.borgbackup.org/ https://www.borgbackup.org/ [2] https://www.stavros.io/posts/holy-grail-backups/ https://www.stavros.io/posts/holy-grail-backups/
- noodlesUK 6y agoDon’t forget restic - it also works great on s3 compatible services! https://restic.net https://restic.net
- nix23 6y agoYes restic is great! No python no dependencies (like borg), just one GO exe and that's it.
- sjellis 6y agoYes, restic would be my pick: easy to deploy and use, well-maintained. If you want some assurance, CERN are deploying restic: https://indico.cern.ch/event/862873/contributions/3724442/ https://indico.cern.ch/event/862873/contributions/3724442/
- josephernest 6y agoI'll have a look too. When using with SFTP, do Restic binaries need to be installed on both local and remote, or just local?
- josephernest 6y agoAs I discover it, borg seems to be pretty good for this use case indeed. But still there is one thing (at least) that can be of interest to some people with nFreezer: it is very simple, it does the job in only 249 lines of code. You can then read the FULL source code in a few hours, to see if you trust it or not. See here: https://github.com/josephernest/nfreezer/blob/master/nfreezer.py https://github.com/josephernest/nfreezer/blob/master/nfreeze... If I want to do this with the source code of the tool you mentioned, I would have to spend at least one full week. (this is normal: this program has 100 times more features). The key point is: if you're looking for a solution for which you don't want to trust a remote server, then you probably don't want to trust the backup tool of a random internet person either. And you probably want to read the source code of the program. So having only < 300 lines of code to read in a single .py can be an advantage.
- giomasce 6y agoI don't see the key point: the two things (trusting the remote storage and the tool) are rather independent for me. I already trust probably billions lines of code which handle my data (basically, every program installed on my system). I don't trust them because I checked all of them, but because I know that the same lines are used by many other people, some of which have actually read parts of them. There is a global trust repository to which every programmer contributes a little bit and every user benefits a little bit. For this reason, I find more trustworthy a project already used and developed by many people, like borg. I didn't just take that for granted, of course: I read the documentation (which, in the case of borg, is very clean and extensive), I looked how it work, I even looked a part of the source code, and I was satisfied. And all of this has nothing to do with where I am putting that data in the end. On the other hand, missing features in the name of small code size is not a great advantage, if those missing features make my backup less reliable or quick or compact. And maybe I end up backing up less stuff because (say) the simple tool does not handle deduplication or compression as effectively as the complicated one, and so is not viable for me. This would be a net loss. Be careful, I am just saying my perspective. You have, of course, the entire right to decide what are your priorities and which tool is better for you.
- rufugee 6y ago
- tonyhb 6y agoAlso want to plug an old HN favourite - tarsnap. https://www.tarsnap.com/ https://www.tarsnap.com/
- jogundas 6y agoThanks for sharing! Is there a reason why borg is excluded from the comparison?
- josephernest 6y agoI will add it. You will probably be surprised, but I wasn't aware of this one when starting my project! In a way, this is fortunate, because it was an interesting journey to code this :)
- quantumofalpha 6y agorestic, borg, even plain rclone. plenty of backup tools exist with encryption at rest, what's the key selling point of yours? I use rclone+restic because of google drive(=cheap unlimited storage) integration in rclone, and restic provides snapshotting on top.
- josephernest 6y agoThanks for asking the question, it's right, there are thousands of backup programs already. The "key thing" that lead me to code my own is that with nearly all the solutions I tried, data was resent over network when a big file is moved to another dir+renamed, WHEN used in encrypted-at-rest mode. It's the case for rsync (the --fuzzy only helps when renamed but stays in same dir), duplicity, and even rclone when in encrypted mode (see https://forum.rclone.org/t/can-not-use-track-renames-when-syncing-to-and-from-crypt-remotes/11063 https://forum.rclone.org/t/can-not-use-track-renames-when-sy...: "Can not use –track-renames when syncing to and from Crypt remotes"). So I wanted to solve this problem, and it works. See point #3 of https://github.com/josephernest/nfreezer#features https://github.com/josephernest/nfreezer#features.
- quantumofalpha 6y ago> The "key thing" that lead me to code my own is that with nearly all the solutions I tried, data was resent over network when a big file is moved to another dir+renamed, WHEN used in encrypted-at-rest mode. restic (and i suppose borg which is similar) solves this problem and goes even further by chunking your files, hashing and deduping the chunks - chunks with the same hash aren't resent. Great e.g. for backing up VM images or encrypted containers where only a small part of the file change - only that small part will be resent between snapshots. Chunking algo is "content-defined", can probabilistically quite efficiently detect shifted chunks and duplicated chunks across different files. (Naturally, all this machinery will also handle the simple cases of renamed and duplicated files on your filesystem)
- bartvk 6y agoNice work. I personally use Arq, which supports SFTP. And like you, I backup to a remote server over which I have no control. Arq is really desktop software, however.
- ajvs 6y agoRclone for encrypted sync (can be virtually mounted) and Duplicacy for encrypted backup. What makes nFreezer different?
- josephernest 6y agonFreezer is different to RClone and Duplicity on a very important point (for me because I often move big multimedia projects): if you move `/path/to/10GB_file` to `/anotherpath/subdir/10GB_file_renamed`, no data will be re-transferred over the network, thus saving 10 GB of data transfer. Indeed, Rclone does not handle renames when in encrypted mode, see https://forum.rclone.org/t/can-not-use-track-renames-when-syncing-to-and-from-crypt-remotes/11063 https://forum.rclone.org/t/can-not-use-track-renames-when-sy... Idem for duplicity (tests done). TL;DR: see 3rd point of https://github.com/josephernest/nfreezer#features https://github.com/josephernest/nfreezer#features ___ Another interesting feature is that you can read/audit the full source code of nFreezer quickly: it's only 249 source lines of code, as of today.
- querez 6y agoI use dar[0] to create incremental, compressed, encrypted backup archives, and then upload those to a GCS bucket. GCS is dirt cheap even for terabytes of data, so as added bonus this gives me a history of all my files (eg, i still have backup archives from 2013). Both calls wrapped into a shell script that is regularly scheduled via a cronjob on my local server. Within my network i sync between laptop, pc and server via syncthing. Works well for me, even though it doesn't handle file renames gracefully. [0] http://dar.linux.free.fr/ http://dar.linux.free.fr/
- josephernest 6y agoDo you think there is a way to avoid to retransfer files that have been renamed/moved with this method? This is the use case I mentioned: I often move files from one dir to another when working on audio- or video-production projects, and it can be dozains of gigabytes.
- deleted 6y ago[deleted]
- querez 6y agoWithin my home network, syncthing handles file renames gracefully [0]. For backups to cloud storage, This might require changes to dar, because according to [1] it's unclear how that tool handles renames. Since I only do these off-site backups to cloud storage once a month, I'm okay with it using a bit more bandwith/storage (as long as it runs over night and doesn't get in anybody's way). [0] https://forum.syncthing.net/t/rename-file/11091 https://forum.syncthing.net/t/rename-file/11091 [1] https://wiki.archlinux.org/index.php/Synchronization_and_backup_programs https://wiki.archlinux.org/index.php/Synchronization_and_bac...
- gradschool 6y ago> How do you work with this? My setup depends only on crappy home made bash scripts and standard utilities, and never leaks encryption keys or filesystem metadata off the local machine. I rsync my home directory with a luks encrypted filesystem image, split the image file into fixed sized chunks using the standard GNU split utility, and then rsync the encrypted split up chunks with a directory on a remote server. The remote server periodically syncs the encrypted image chunks with S3 using the usual aws cli tools, and runs a script to "cat" the chunks back into a remote copy of the whole image and display the overall md5 sum. Another machine off site keeps amazon honest by syncing the s3 bucket with a local copy, reassembling the encrypted filesystem image, and computing the md5 sum. An extra benefit is that I can use the filesystem image on the remote server over sshfs as the backing device for a locally loopback mounted luks partition to access or modify individual files without downloading the whole image or corrupting it, provided no more than one session is open at a time. Back when I was more paranoid, I used only a raw dmcrypt image rather than a luks encrypted image in this setup so that there would be no luks header to give it away and I could plausibly claim that it was one-time-pad encrypted data, for which I could construct a key to yield any plaintext of my choice. (Yes, I know I won't be such a smartass when it comes to a rubber hose attack.)
- wheybags 6y agoPersonally, I use zfs on my home server. All my other machines rsync to a backups folder on there, and once a month a scheduled email reminds me to grab my backup usb hard drive and do a zfs send | recv to sync for an offline backup. If you were using an untrusted server as a zfs send target, as of a few months ago, you can actually do an encrypted send to a remote without unlocking on the remote, which is pretty cool.
- DarkmSparks 6y ago>Being curious: how do you usually do remote backups of your important files? I keep my important files in a ~50GB veracrypt container then backup that to multiple locations including a local usb disk every day or so with a shell script. dropbox nuked its copy once but it worked fine again after a name change (version number)
- basicneo 6y agoDo you copy the entire 50GB on every backup, even if you've only changed 1MB?
- DarkmSparks 6y agoTo an external 1TB HDD yes. Keep one for the end of each month. Veracrypt containers rsync and dropbox quite nicely, just keep one copy of those.