4 ms·
> I guess the question is whether rsync is using multiple threads or otherwise accessing the filesystem in parallel No, that is not the question. Even Wikipedi
by nh2 8mo ago
> I guess the question is whether rsync is using multiple threads or otherwise accessing the filesystem in parallel
No, that is not the question. Even Wikipedia explains that rsync is single-threaded. And even if it was multithreaded "or otherwise" used concurent file IO:
The question is whether rsync _transmission_ is pipelined or not, meaning: Does it wait for 1 file to be transferred and acknowledged before sending the data of the next?
Somebody has to go check that.
If yes: Then parallel filesystem access won't matter, because a network roundtrip has brutally higher latency than reading data sequentially of an SSD.
- dekhn 8mo agoNote that rsync on many small files is slow even within the same machine (across two physical devices), suggesting that the network roundtrip latency is not the major contributor.
- nh2 8mo agoThe original post only mentions 3564 files and rsync spending 8 minutes on that. This just doesn't check out.
- Dylan16807 8mo agoThe filesystem access and general threading is the question because transmission is pipelined and not a thing "somebody has to go check". You just quoted the documentation for it. The dead time isn't waiting for network trips between files, it's parts of the program that sometimes can't keep up with the network.
- nh2 8mo agoI quoted the documentation that claims _something_ is pipelined. That is extremely vague on what that is and I also didn't check that it's true. Both the original claim "the issue is the serialization of operations" and the counter-claim all sound like extreme guesswork or me. If you know for certain, please link the relevant code. Otherwise somebody needs to go check what it actually does; everything else is just speculating "oh surely it's the files" and then people remember stuff that might just be plain wrong.
- Dylan16807 8mo agoSpeculation isn't the most useful thing, but saying "that is not the question" to valid speculation is even less useful.
- nh2 8mo agoI was about to say: The question was what exactly rsync pipelines, and whether it serialises its network sends. If true, that would be a plausible cause of parallelism speeding it up. Serial local reads are not a plausible cause, because the autor describes working on NVMe SSDs which have so low latency that they cannot explain that reading 59 GB across 3000 files take 8 minutes. However: You might actually be half-right because in the main output shown in the blog post, the author is NOT using local SSDs. The invocation is `rsync ... /Volumes/mercury/* /Volumes/...` where `mercury` is a network share mount (and it is unspecified what kind of share that is). So in that case, every read that looks "local" to rsync is actually a network access. It is totally possible that rsync treats local reads as fast and thus they are not pipelined. In fact, it is even highly likely that rsync will not / cannot pipeline reading files that appear local to it, because normal POSIX file IO does not really offer any ways to non-blocking read regular files, so the only way to do that is with threads, which rsync doesn't use. (Extra evidence about rsync using normal blocking writes, and not supporting threads, beyond the fact that no threading code exists in rsync's repo: https://github.com/RsyncProject/rsync/blob/236417cf354220669014317b1ba818b9d931afbb/rsync3.txt#L365 https://github.com/RsyncProject/rsync/blob/236417cf354220669...) So while "the dead time isn't waiting for network trips between files" would be wrong -- it absolutely would wait for network trips between files -- your "filesystem access and general threading is the question" would be spot-on. So in that case rclone is just faster because it reads from his network mount in parallel. This would also explain why he reports `tar` as not being faster, because that, too, reads files serially from the network mount. Supposedly this situation could be avoided by running rsync "normally" via SSH, so that file reads are actually fast on the remote side. The situation is extra confused by the author writing below his run output: even experimenting with running the rsync daemon instead of SSH when in fact the output above didn't rsync over SSH at all. Another weird thing I spotted is that the rsync output shown in the post Unmatched data: 62947785101 B seems impossible: The string "Unmatched data" doesn't seem to exist in the rsync source code, and hasn't since 1996. So it is unclear to me what version of rsync used. I commented that on https://www.jeffgeerling.com/blog/2025/4x-faster-network-file-sync-rclone-vs-rsync/#remark42__comment-df15b89b-baf6-47b0-a7e9-c6e6553d1ff3 https://www.jeffgeerling.com/blog/2025/4x-faster-network-fil...