10 ms·
The fastest rm command and one of the fastest cp commands
- guiand 4y agoThis could be useful for e.g. bazel, where I’ve regularly seen deleting the bazel caches take on the order of 10 minutes because of absurdly large and repeated directory trees (caused by things like runfile trees containing the python interpreter).
- eru 4y agoThe lazy delete as outlined in https://news.ycombinator.com/item?id=35309328 https://news.ycombinator.com/item?id=35309328 should also work?
- rektide 4y agoI cant wait to see the future where we have uring (io_uring) based tools that all work async, getting hopefully embarassingly parallel. Heck yes copying & deleting 2x as fast. Actually I'm on btrfs so reflink copy is like instant I think? I should test that better.
- zamalek 4y agoasync nearly always trades concurrency for latency, so it wouldn't necessarily be faster overall. Lots of profiling and tuning would probably be involved. That, or your algorithm could be optimal for concurrency already - and you would see an immediate performance improvement.
- vlovich123 4y agoWhat? How does async (and io_uring specifically) give up concurrency? They’re orthogonal things.
- zamalek 4y agoYou interpreted it the wrong way. You gain concurrency, you gain latency. With blocking I/O and parallelism you have a thread ready to go when the operation is complete. You have N threads for N iops. With concurrency you have to dequeue completed work, and then delegate that work to (usually) fewer than N threads. Dequeuing completed iops takes time (it's an extra syscall), and there may not be a thread ready hand the completed iop. More latency. Running 1000s of threads isn't realistic because your OS would typically grind to a halt, so concurrency is unavoidable. It does have a cost, though.
- vlovich123 4y agoI’m not aware of any latency impact from io_uring. If anything, it has lower latency because you can pipeline I/O from a single CPU which you can’t do from typically thread-based parallelism. Additionally, the processing of the ring buffer could happen on a background kernel thread (in theory not sure if it happens today) which then avoids context switching away from your thread and screwing with cache performance.
- zamalek 4y agoio_uring avoids the syscall overhead. Consider this intentionally bad scenario: you have one background thread pulling completions off the ringbuffer, and passing those completions to a single worker thread (nodejs would be a real example of this). In this scenario the latency of the first completion would be fantastic, but you'd have to wait for the single to become available for subsequent completions. To be clear, the added latency here is better than the work never happening at all (which would be the result of running 1000s of threads on modern mainstream operating systems), but there is unavoidable latency if you are handling >N iops with N threads (which is intrinsic to the definition of concurrency). I am referring to the broad, general case, much like big-O works. You can find numerous exception to big-O, such as preferring arrays over hashes when the set is very small. Let's invent big-L notation, N is the number of threads, M is number of iops. With pure parallelism you have L(N), with pure concurrency you have L(M), and with a hybrid you have L(M-N).
- 4y ago
- SUPERCILEX 4y agoWe're working on this! https://github.com/axboe/liburing/issues/830 https://github.com/axboe/liburing/issues/830
- deleted 4y ago[deleted]
- tedunangst 4y agoIf you need a large sub tree gone, mv dir .old-dir && rm -r .old-dir & works pretty well. Faster even than an optimized parallel unlinker.
- bee_rider 4y agoThat’s pretty outside-the-box and fun.
- eyelidlessness 4y agoAmusingly it’s basically how most GUI desktops “delete” stuff, they just put a confirmation in place of the &&
- HeavyFeather 4y agoI don't think any GUI move items before a confirmation, so I don't see the link at all.
- eyelidlessness 4y agoThey almost universally do! They move stuff to a designated trash or recycle bin or whatever to stage for final deletion when you commit to it.
- Dylan16807 4y agoWindows, if asking is enabled, will ask before moving into the recycle bin. Once the file is in the recycle bin, it will probably be months before final deletion happens, and windows will not ask before doing so.
- ElectricalUnion 4y agoUnless you happen to start to run low on disk space; at that point, Storage Sense will kick in and start complaining about it - if you enable it without changing standard settings it deletes "trash bin" files every 30d.
- 8organicbits 4y agoThe benchmarks are impressive: https://github.com/SUPERCILEX/fuc/tree/master/comparisons#results https://github.com/SUPERCILEX/fuc/tree/master/comparisons#re...
- saagarjha 4y agoHmm. I appreciate the cute name for the project but I tested this on my machine (a Mac) and the results were not very impressive. I sacrificed a few of my SSD cycles to test how this is on deleting Xcode (for those unaware, it's a 11 GB mess of several hundred thousand files of varying sizes). Here are the results: $ time rm -rf Xcode.app real 0m39.850s user 0m0.429s sys 0m29.153s $ time rmz Xcode.app real 0m36.476s user 0m1.468s sys 1m59.916s It's a little bit faster, but not by much. Despite claims that it runs in parallel, it seems like it really just hits unlink on many more cores and contends on the kernel's filesystem spinlocks rather than doing useful work. So the end result is that rm uses 70% CPU and rmz uses 400% CPU and they basically end up doing the same thing. (FWIW, I don't use rm when deleting Xcode anyways, because it takes long and I do it too often. When testing unxip I just have it write to a temporary APFS volume and wipe it in between runs, which takes all of 10 seconds.)
- eru 4y agoSlight tangent: why do you need to delete Xcode so often? (I don't really develop on Mac, mostly on Linux.)
- oefrha 4y agoThey develop https://github.com/saagarjha/unxip https://github.com/saagarjha/unxip, “a fast Xcode unarchiver”. Very few people should routinely delete Xcode.
- eru 4y agoThanks!
- LeoPanthera 4y agoNot OP, but disk space maybe? It's pretty big. $ du -Ash Xcode.app 22G Xcode.app
- saagarjha 4y agoXcode is compressed on disk by default; you should generally use du -sh when measuring its size.
- preseinger 4y ago> The key insight is that file operations in separate directories don’t (for the most part) interfere with each other, enabling parallel execution. i'm clearly missing something here parallel execution helps when operations are cpu bound file operations are (almost always) io bound and totally unclear how directories represent an "interference" boundary bizarre
- charcircuit 4y agoWhat I would suppose is that it reduces the amount of redundant IO where a directory is edited, flushed to disk, and then later updated again. If all these updates happen in a batch there will be less IO overall.
- preseinger 4y agosure, but the fs cache does all of this kind of stuff for you, right?
- mort96 4y agoParallel execution absolutely helps when operations are IO bound, if they're more or less independent. Making two network requests in parallel is twice as fast as making them sequentially, if the payload is small enough so that latency dominates and bandwidth is negligible. The question is, how independent are IO operations in separate directories. And the article is claiming that they're fairly independent and don't block each other.
- preseinger 4y agomaking two io-bound requests in parallel is twice as fast as making them sequentially, only if they don't contend for the same io resource -- bandwidth, disk iops, etc. maybe this is what you mean by independent? but the thing is that in disk io, directory structure is (as far as i know) basically unrelated to relevant contentious resources, when measuring speed maybe if you're doing a billion small files than overhead begins to matter, but copying 3 big files from 3 different directories is gonna take just as long if you do them in parallel vs. if you do them sequentially that may not be true if they're on different disks, but that kind of proves my point, the directory isn't the factor, the underlying disk is > The question is, how independent are IO operations in separate directories. And the article is claiming that they're fairly independent and don't block each other. yeah and in this sense the article is misleading, because (as far as i know) directories are basically unrelated to independence in the general case
- jeffrallen 4y agoI once wrote a command called trickle-rm, which was designed to be an i/o constrained rm, this the exact opposite of this article. I needed trickle-rm because after extensive analysis, I'd found that the i/o load of a diagnostic data cleanup job was interfering with the very tight latency requirements of the main app on the server. My first effort was "nice rm -rf $oldLogs". When I'm feeling a bit evil I will ask an interview candidate why that didn't work. Always something interesting to learn at the margins
- mherdeg 4y agoWhy didn't it work? Was try 2 ionice?
- b112 4y agoNot the parent, but ionice only works with certain elevators. And elevator choice mattered a lot on spinning disks.
- jeffrallen 4y agoNice reduces your scheduling priority for access to the CPU, which means it might take your niced rm longer to do the processing to submit a bunch of i/o to the kernel. But once those i/o ops are in the kernel, they are competing directly on an equal playing field with the latency critical i/o ops of the real app, causing the same degradation of latency to the end users. ionice was not available on my platform, thus the invention of trickle-rm. Now the second interview question: how did trickle-rm work? How can you simply "reflect" the back pressure of the "x rm's per second" constraint back onto the "walk the directory tree to find the rm's to do" so that the tree walk generates a trickle of i/o operations?
- aeonik 4y agoRun rm in a bash loop for each file and sync after each iteration? I'm surprised that rm's io would cause that much of an issue. I thought it only removed entries from the partition table.
- gambiting 4y agoOn windows I find it's much faster to use robocopy to mirror an empty folder into a full one than it is to delete the full folder than pretty much any other command. I regularly have to delete a 200GB folder with 500k+ files in it, and robocopy outperforms regular rm -r by a factor of two, and GUI shift+right click delete by a factor of 4 or 5.
- nomercy400 4y agoWhen deleting a 25k-file node_modules folder on windows, I use the built-in rmdir, as it is much faster than the GUI. I do not know how it compares to robocopy.
- criddell 4y agoI’m going to have to try that. I want to throw my computer out the window when I delete a folder in File Explorer and I get the “Preparing to delete” dialog come up. If I’m truly deleting and not recycling, what is it preparing?
- gambiting 4y agoIt queries the list of all files and their sizes so it can show you a progress bar. RM and robocopy will just get on with it, but they don't know how much they are about to delete(not that they need to)
- shaggie76 4y agoThat's a neat trick; I was in a similar situation recently and wrote a tool that's faster for me though; for a million files ROBOCOPY was 257 seconds, but https://github.com/shaggie76/FastDelete https://github.com/shaggie76/FastDelete did it in 34 seconds on a hex-core laptop.
- gambiting 4y agoOh that's super cool! Will need to try it.
- redman25 4y agoJust started working on ‘can’, an ‘rm’ replacement, that moves files to the trash instead instead of deleting them. It only works on macos right now but I was hoping to make it cross-platform. It’s not faster than ‘rm’ but hopefully saves accidental deletions. https://github.com/joshvoigts/can https://github.com/joshvoigts/can
- plg94 4y agoFor Linux there's [trash-cli](https://github.com/andreafrancia/trash-cli/ https://github.com/andreafrancia/trash-cli/). Doesn't seem to work for MacOS per this issue (https://github.com/andreafrancia/trash-cli/issues/284 https://github.com/andreafrancia/trash-cli/issues/284), but it suggests to use https://hasseg.org/trash/ https://hasseg.org/trash/
- redman25 4y agoYa, it’s mostly a learning excuse.