5 ms·
I'm somewhat of an old fart, but rsync(1) has worked well for me. It integrates well with ssh on both sides, using ssh channels to execute the binary on the ser
by perbu 5y ago
I'm somewhat of an old fart, but rsync(1) has worked well for me. It integrates well with ssh on both sides, using ssh channels to execute the binary on the server side and transporting data nicely back and forth.
It isn't suited for millions of files, but neither is scp.
- gerdesj 5y agoIt handles as many files as I've ever thrown at it - often in the millions. The great thing about rsync is it is restart-able and the source and destination don't even have to be local.
- perbu 5y agoIt handles millions, but it can be a lot faster to just pipe output from tar through the ssh connection.
- hotpotamus 5y agoGot any benchmarks/write ups on the subject? I did a bit of testing myself a long time ago and basically the answer just ended up being to use rsync because any differences were marginal. That said, I didn't test with millions of files.
- perbu 5y agoI think perhaps this was a bigger issue back in the day when we were using rotating harddisks. In those day doing a seek would be a lot slower than doing a write. Today seeks are mostly instants, so maybe my experience isn't valid anymore.
- porker 5y agoYour experience is valid today but for a different reason: if you're comparing millions of tiny files there's a lot of back and forth. If you're streaming a single archive, it only checks if that single file has been modified. Like everyone here I've no benchmarks but have got burned trying to rsync around too many small files.
- nani_o 5y agoI got into the « lot of small files » case, and I ended up generating a file list, split it and feed multiple rsync instances with xargs.
- paulmd 5y agothe biggest weakness of rsync is that it's single-threaded. Single-thread-copy with millions of small files is painfully slow, Rsync is as good as it can be but you just need more threads. I think that's what a sibling is getting into with the "better to tar files and send them over ssh in some cases" thing. And yes, you can hack it in after-the-fact with xargs/etc but it's clunky compared to just having native multithreading like rclone/etc.
- jimpudar 5y agoIf you use GNU Parallel, there's a fairly easy way to parallelize [1]. That being said, I hadn't heard of rclone - thanks for mentioning it, it looks amazing. I'll definitely be trying this out for my use cases... [1] http://www.gnu.org/software/parallel/man.html#example-parallelizing-rsync http://www.gnu.org/software/parallel/man.html#example-parall...
- robertlagrant 5y agoWow, Parallel also looks awesome. Never seen it before.
- SEJeff 5y agoWe've put rsync in a HPC scheduler and used it, or some tooling ontop of it really, to copy billions of files for a large-ish HPC compute cluster with many P of data.
- gerdesj 5y agoThat's my experience - millions or billions of files - who cares? rsync is time served. It just works. I'm sure there are other funky solutions but they are not proven over decades.
- johnthescott 5y agowe have had problems with millions of rsunk files.
- photon-torpedo 5y agoAlso, rsync can create more faithful copies than scp. E.g. scp can't copy a symlink as a symlink, instead it will follow the symlink and copy the file or directory.
- FullyFunctional 5y agoThis is one of the main reasons I never use scp and always use rsync -a. IMO, seems a poor design of scp.