3 ms·
My preferred solution is rmlint [https://github.com/sahib/rmlint https://github.com/sahib/rmlint] mostly because it also looks at duplicate directories. It prod
by tryptophan 3y ago
My preferred solution is rmlint [https://github.com/sahib/rmlint https://github.com/sahib/rmlint] mostly because it also looks at duplicate directories. It produces a bash script instead of deleting anything itself, so you can examine it before running the script it made.
- jbaber 3y agoThanks for the recommendation. I find unleashing fdupes on my precious precious data too terrifying
- runlevel1 3y agormlint has a neat feature where it can write the checksum to the file's xattr (along with the mtime when it's calculated) so that it doesn't need to recalculate it in future passes. If the file comes up as a dupe candidate again (when its size matches another file's) and if its current mtime matches the one in the xattr, then it will use the hash from the checksum. In practice, I've only had this break a couple times. In both cases, it was because it scanned an incomplete download from a tool configured to set mtime to the server's Last-Modified date. When the download was later finished, the downloader backdated the mtime again. So the mtime matches but the content no longer matches the checksum. So you end up with a false negative on the completed file until that attribute is cleared. Which I'm ok with.