4 ms·
The solution is to upgrade to rsync.
by DarylZero 5y ago
The solution is to upgrade to rsync.
- erdewit 5y agoAnother great option is to mount the remote file system with SSHFS.
- formerly_proven 5y agosshfs/sftp throughput is very limited. scp acts much better here.
- egberts1 5y agorsync would be the fastest of the lot, plus it does file permissions as well.
- acdha 5y agosftp is often faster since it starts copying immediately and hides latency more effectively and both sftp and scp copy permissions if you ask them to: -p Preserves modification times, access times, and modes from the original files transferred.
- remram 5y agoWouldn't `rsync --ignore-existing` have similar latency from sftp then? It will be slower than rsync without the option, but not slower than sftp.
- suifbwish 5y agorsync accepts pipes which means you can run end to end compression and decompression automatically during the transfer. Processes can also be parallelized. Most of the time when I see someone claiming rsync does not suit their purpose it ends up being that they aren’t confident in using it safely. Most of the time I never use anything else except the -av flag
- rwmj 5y agoIt's more likely we'll bite the bullet and use sftp directly from libssh. rsync doesn't work for our use case (and nor does sshfs).
- tialaramex 5y agoSFTP makes sense here, that's why OpenSSH is deprecating scp because SFTP is much better defined. I actually wrote all my scp-style scripting with SFTP for years, so much so that I think projects I did this for lived their entire lifecycle and are gone (but I don't know anybody who still works at the place I wrote the earlier ones) without running into the problem I feared - it's just not clear what scp means for non-trivial cases.
- loudtieblahblah 5y agoHate to be that guy making you defend a situation you know more about than any of us here.. But I incredulously am curious as to why rsync wouldn't work for you.
- acdha 5y agoI can't speak for them but there are two big ones I've hit: 1. rsync likes to build the entire list of operations before doing anything. If you have workloads which requires scanning lots of small files and/or using something like NFS / s3fs, etc. that means a LOT of I/O delays before a single byte of payload is transferred. 2. The checksum algorithm is really cool if you have large files which only change a little but it's relatively expensive and brittle. I've run into multiple cases where naive sftp was faster because we had a high-bandwidth, high-latency network connection.
- quesera 5y agoFWIW, 1. This is no longer true in rsync-3.0.0 (Mar 2008) and later. It does build the list of course, but transfer begins after establishing just a few directories of content. This "incremental" mode is not available if you specify an option that requires the full list to begin, documented under the "--recursive" option. 2. Checksums are not computed by default. If the files match time/size, they will not be transferred or checksummed unless "--checksum" is turned on. All files which are transferred, are compared by checksum when complete but this is not meaningfully expensive since the IO is already done. Your issues with high-bandwidth, (very) high-latency links make sense. The rsync algo was designed to minimize bytes-sent over the undersea cable to Australia (low-bandwidth, moderate-latency). This still works well for most internet traffic, but not if your latency numbers are way out of balance!
- LinuxBender 5y agoTo add to this for completeness sake, if the account is sftp-only one can use the lftp client with its mirror subsystem that replicates most of the functionality of rsync.