6 ms·
Casync – A Content-Addressable Data Synchronization Tool
- xyzzy_plugh 4y agoMentioning the more portable desync is obligatory: https://github.com/folbricht/desync https://github.com/folbricht/desync
- ranit 4y agoHow is go more portable than C?. Genuine question.
- ReactiveJelly 4y agoI think they mean to OSes that don't run systemd / Linux kernel?
- deleted 4y ago[deleted]
- ranit 4y agoIt is not systemd dependent.
- jeremyjh 4y agoThe repo linked answers this question.
- beagle3 4y agodesync is what git-lfs should have been: rolling hash chunk based storage, with local caching of chunks, proxying of remote chunks, etc. When I have big 100MB binary files that i want to version, the changes are small (1MB in one place, and a few more KB in others). I also have a few multi GB SQLite databases I would like to version where this would help. (Changes are less than 5% of the data, but index pages throughout the file mean rsync-style partitions still transfer 50% or so of the file. To actually achieve storage efficiency, I store textual dumps, and they also fit better with borg/restic/casync than git or git-lfs Would have been extra awesome if desync would be able to use a git repo as storage.
- pabs3 4y agogit itself should have been rolling hash chunk based storage. Storing very large text files (like bash history or the Debian security team git repo) in git is very inefficient due to the current design.
- beagle3 4y agoIt is inefficient up until the point where you ask git to do a gc+repack, at which point git will look for deltas - and if it is changes to the same file, it will likely find them and encode them as a delta. For text files with really small changes, this is comparable to a "diff" size, and is better than what a rolling hash would achieve. However, for text files with larger changes, and of course for binary files - a rolling hash is much more effective. Additionally, a rolling hash would easily reuse cross-file similarity, where as the current delta-finding code is likely to only find redundancy in the history of the same file (or similarly named in the same directory).
- ReactiveJelly 4y agoWhy is this under systemd? I hope I don't have to update all of systemd to get updates to a sync tool.
- FredFS456 4y agoFrom the announcement blog post (http://0pointer.net/blog/casync-a-tool-for-distributing-file-system-images.html http://0pointer.net/blog/casync-a-tool-for-distributing-file...): > Is this a systemd project? — casync is hosted under the github systemd umbrella, and the projects share the same coding style. However, the code-bases are distinct and without interdependencies, and casync works fine both on systemd systems and systems without it.
- denton-scratch 4y ago> However, the code-bases are distinct and without interdependencies Didn't they say that about udev?
- deleted 4y ago[deleted]
- viraptor 4y agoI was wondering how this gets any common chunks at all with the removed file boundaries. Turns out that chunks don't have a set size, just min/max/avg values, so unaligned streams may end up synchronizing. https://github.com/systemd/casync/blob/master/src/cachunker.c https://github.com/systemd/casync/blob/master/src/cachunker.... If I understood that correctly, that's pretty cool. But looking at the code I'm having strong "nope" feelings. First, because of lines like "q += m, n -= m;". Second, because of int/enum/semantic abuse: `compression_type` may be _CA_COMPRESSION_TYPE_INVALID which I hope is -1, `>= 0` as a known compression type, or `-EAGAIN` as an error. (from https://github.com/systemd/casync/blob/99559cd1d8cea69b30022261b5ed0b8021415654/src/cachunk.c#L55 https://github.com/systemd/casync/blob/99559cd1d8cea69b30022... ) I'd bet that just throwing afl at the decompressor will find issues :( (source - I had this feeling about systemd-resolved and threw afl at it, found issues) I do like the idea though.
- klysm 4y agoHow hard is it to set up AFL? I’ve never had the opportunity to set up a fuzzer
- viraptor 4y agoDepends on the app / file format. If you can isolate some trivial behaviour on a single file you may not even need to change the app. For example "casync list --store=/var/lib/backup.castr input.caidx" could maybe be used to fuzz the index reading if it doesn't care about looking into the store immediately. And if it does, you could comment out those parts as needed. You can get very fancy with the process, but you can start with "here's your single valid input sample, go!" Check out https://fuzzing-project.org/tutorial3.html https://fuzzing-project.org/tutorial3.html and the AFL guide.
- klysm 4y agoThanks! I'll be looking at using that for some of our code
- SaveTheRbtz 4y agoI was recently playing[0] with the ZSTD seekable format[1]: a *valid* ZSTD stream format (utilizing ZSTD skippable frames, that are ignored by de-compressor) w/ an additional index at the end of the file that allows for random access to the underlying compressed data based on uncompressed offsets. This combined w/ a Content-Defined-Chunking[2] and a cryptographic hash[3] in the index allows for a very efficient content-addressable storage. For example I've successfully applied it to a bazel-cache, which gave me between 10x and 100x wins on repository size w/ negligible CPU usage increase. [0] https://github.com/SaveTheRbtz/zstd-seekable-format-go https://github.com/SaveTheRbtz/zstd-seekable-format-go [1] https://github.com/facebook/zstd/blob/dev/contrib/seekable_format/zstd_seekable_compression_format.md https://github.com/facebook/zstd/blob/dev/contrib/seekable_f... [2] e.g FastCDC https://www.usenix.org/system/files/conference/atc16/atc16-paper-xia.pdf https://www.usenix.org/system/files/conference/atc16/atc16-p... [3] https://github.com/facebook/zstd/pull/2737 https://github.com/facebook/zstd/pull/2737
- kldx 4y agoIs your bazel cache implementation open source? I am dabbling in bazel and I am not sure where zstd fits in the bazel cache model. I'm interested in learning more about this
- SaveTheRbtz 4y agoI did PoC experiments with compression, chunking, and IPFS here: https://github.com/SaveTheRbtz/bazel-cache https://github.com/SaveTheRbtz/bazel-cache If you need a mature compression implementation for bazel I would recommend using recent bazel versions w/ gRPC-based bazel-remote: https://github.com/buchgr/bazel-remote https://github.com/buchgr/bazel-remote bazel nowadays supports end-to-end compression w/ `--experimental_remote_cache_compression`: https://github.com/bazelbuild/bazel/pull/14041 https://github.com/bazelbuild/bazel/pull/14041
- rwmj 4y agoReally wish this was part of official zstd (https://github.com/facebook/zstd/issues/395#issuecomment-535875379 https://github.com/facebook/zstd/issues/395#issuecomment-535...) and not a contrib / separate tool.
- raggi 4y agoWe did all this stuff for packages in fuchsia and it works really well. We went further though, the on-disk storage is a write-once immutable store of read-verified data.
- dboreham 4y agoSeems quite similar to IPFS, but no comparison in the readme/article.
- joelg 4y agoIt's much more similar to a related Protocol Labs project: CAR (content-addressable archives) https://ipld.io/specs/transport/car/ https://ipld.io/specs/transport/car/