5 ms·
I wonder how it compares to restic or borg. Besides the gui anyway…
by gingerlime 5y ago
I wonder how it compares to restic or borg. Besides the gui anyway…
- benrockwood 5y agoLooks like a polished restic to me.
- isbvhodnvemrwvn 5y agoPolished*
- StavrosK 5y agoA hard joke to get, but I liked it.
- Scaevolus 5y agoRestic doesn't support compressing backups, and kopia does. Otherwise, the architectures appear to be very similar.
- lucb1e 5y agoHow much space do you typically save by compressing these days? Given that even smallish things like documents are already compressed archives, pictures/audio/movies of course already have heavy purpose-specific compression. The main things I can still think of that are sparse on purpose are database files and disk images (not very mainstream, but also not uncommon). So like, a few gigabytes per terabyte (a few promille) unless you're really heavy on either databases or virtual machines? I can see why one would like to enable it, but deduplication (which breaks if you naively implement compression, iirc that's why restic hasn't yet implemented it) is much more worth it because it enables incremental backups and you don't, for example, have to worry about making a copy of another system that has many of the same files (think game files or system files).
- wmf 5y agoI assume compression works well for source code and other developer artifacts. Obviously you have to do dedupe, then compression, then encryption.
- lucb1e 5y agoSource code compresses very well indeed, but my hunch is that it's peanuts. Let's see, I've got a projects directory with various projects from the past decade (all custom, there's a separate dir for downloaded repositories). I've mostly written things in Python and PHP (the JS/CSS/HTML stuff is on a server mixed with things like owncloud or SMF or so; harder to isolate). PHP: 591 KiB, 196 files, 13'813 LOC, 1'236 comment lines. Python: 345 KiB, 136 files, 7'760 LOC, 930 comment lines. If someone spends 5 minutes of developer time trying to compress that to save disk space, that's already not worth it. Also in huge projects, the actual code is not going to be taking gigabytes of space. And if you mean in git history: that is, again, already compressed. Other developer artifacts: I've got 26 GiB of project directories, this time including downloaded software and it will also include binaries (hashcat and jtr are in there, I wouldn't be surprised if there's also a medium-sized dictionary or two). Doing tar c . does not seem to add much overhead (26.5 GiB). Compressing that stream with pigz -1 (multithreaded gzip) brings it down to 17 GiB. 35% off is better than I thought! I wonder which files compress so well, hmm let me `find -type f | shuf | head -9001 | while read line; do echo "$(($(wc -c <"$line")-$(<"$line" pigz -1 | wc -c))) $line"; done | sort -n`... The largest difference is a huge 121 MiB binary that compresses down to 36 MiB. I didn't know these files were so sparse (not a C(++) dev), interesting! While I'm looking into this, let's also look at my "documents" directory. It's 43 GiB and compresses down to 33 GiB. Not as good, but still worth it, more than I thought! (And this compression isn't good, but probably not more than 10% worse before the compression gets impractically slow.) It might not quite get a total backup size down by a disk size (e.g. not 1T down to 500G), but it definitely allows to keep more history before having to worry about what you want to keep and what you want to toss.
- aidenn0 5y agoFor local backups, compression is less of an issue; for things that are compressible, transparent file system compression seems to get about 110% the size of what I would get by any non-CPU bound levels of compression using a tgz. Since (as others in this thread have noted), compressible files tend to also be smaller files (the only exception I can think of would be if your log rotation doesn't compress old logs), the fact that only a fraction of what I backup is 10% larger is kind of "okay." When you're sending across the network though it can be a big deal.
- MikusR 5y agoOn disk the backups are encrypted, that means no transparent compression.
- isbvhodnvemrwvn 5y agoFor people who haven't dealt with this - a good encryption scheme produces output which you can't tell apart from a purely random stream of bits - it has very high entropy, and is therefore not compressible.
- zmix 5y agoThanks for clarifying this. That must mean then: first encrypt, then compress?
- aidenn0 5y agoCompress than encrypt.
- isbvhodnvemrwvn 5y agoAnd this has some pitfalls as well, ciphertext in general has about the same length as the plaintext, so compression rate can be used to infer some information about the plaintext. It's more of a problem with interactive protocols than with backups, but still worth keeping in mind.
- poronski 5y agoGotta say kopia does have a very strong restic vibe to it. Not a bad thing, just means that restic managed to get lots of things right.
- ntolia 5y agoYou can see some of the performance differences here - https://blog.kasten.io/benchmarking-kopia-architecture-scale-and-performance https://blog.kasten.io/benchmarking-kopia-architecture-scale...
- jeremyw 5y agoNote this compares an older restic version that doesn't include the order of magnitude improvements in cloud communication.