4 ms·
What about "comm" - compare two sorted files line by line. You can easily get occurrences only in file 1, in both files, only in file 2. Super powerful and sav
by nunoferreira 7y ago
What about "comm" - compare two sorted files line by line.
You can easily get occurrences only in file 1, in both files, only in file 2.
Super powerful and saved me hours of work.
- pimlottc 7y agocomm is a really useful tool, with one big caveat — you must make sure your input files are all sorted the exact same way. If not, you can get unexpected results, and worse, might not even realize it. This may seem obvious, but there are many tiny ways that sorts can differ between locales, operating systems and programs (e.g. Excel), especially when dealing with Unicode. It may look the same 99% of the time, and you may not realize until later that you’ve accidentally filtered out values.
- nunoferreira 7y agoAbsolutely! From my experience I only use with listings from the same source with the same sort tool (mostly unix sort).
- TomNomNom 7y agoMy advice is to sort the files just-in-time using the shell: comm <(sort fileA.txt) <(sort fileB.txt)
- Hello71 7y agoGNU comm prints a warning if either file is not sorted, unless all input lines are pairable.
- zamadatix 7y agoComm is perfect for scripting usage but you might find diff better for human usage. Added bonus diff also does binary. Plus diff was in part written by the author of the linked content :).
- TheGrassyKnoll 7y agoYou might enjoy tkdiff sudo apt-get install tkdiff
- loeg 7y agoComm operates on sets. Diff is a patch generator. They serve different needs. They're both useful!