4 ms·
3) if all the lines of the 10+GB file are actually unique, wouldn't awk keep the whole file in RAM? For files larger than my RAM could this leave my system unre
by JoeAcchino 14y ago
3) if all the lines of the 10+GB file are actually unique, wouldn't awk keep the whole file in RAM? For files larger than my RAM could this leave my system unresponsive because it's thrashing on swap?
- niggler 14y agoThe sort | uniq method literally needs to sort the file and pipe it to uniq, a far more memory-intensive operation than the single-pass awk check. You can write your own hash function in AWK if you think you may overstep memory, but of course you risk hash collisions. It's a tradeoff. I tried it on a 1U server with 24GB ram a few years ago and found that the sort was thrashing at the 10GB file size while AWK handled it easily
- carlesfe 14y agoYes, awk does some real black bagic. It's awesome how it can parse really, really big files.
- stevvooe 14y agoSort will actually externally sort blocks into temp files and merge them. Adjusting this block size can help with thrashing. Awk may still be better for uniquification.
- enigmo 14y agoWhat about 'sort -u' ?
- stevvooe 14y agoI'm not sure. I haven't closely studied the difference between each algorithm. My guess would be that sort -u would perform better as the data set gets larger with a good block size setting because it does do an external sort. Cardinality would also affect the performance. If the unique set handily fits in memory, an external sort on a large data set wouldn't be very efficient.
- gav 14y agoOr you can use `sort -u` and not have to pipe to `uniq`.
- alexkus 14y agoYes it will, and a bit (read: a lot) more than 10GB as it needs to store the contents of variable x in a hash table (with the corresponding hash key and value of the counter). There's no other magic way it can 'know' whether a particular line has been seen before. You can't rely on hash keys alone as the hashes aren't guaranteed to be unique. For files with relatively few duplicates it's going to be a lot slower than sort | uniq. Trying it on a 128MB file (nowhere near enough time to test a 10GB file) filled with lines of 7 random upper case characters[1] (so hardly any duplicates):- $ wc -l x.out 16777216 x.out $ time ( sort x.out | uniq ) | wc -l 16759719 real 0m17.982s user 0m42.575s sys 0m0.876s $ time ( sort -u x.out ) | wc -l 16759719 real 0m20.582s user 0m43.775s sys 0m0.688s Not much difference between "sort | uniq" and "sort -u". As for the awk method:- $ time awk '!x[$0]++' x.out | wc -l has been running for more than 20 minutes and still hasn't returned. For that 128MB file the awk process is also using 650MB of memory (according to ps). Will check up on it later (have to go out now). This Linux machine has ~16GB of memory so the file was going to be completed cached in memory before the first test. All things considered equal the awk method will be roughly O(n) (e.g. linear against file size) and sort/uniq will be O(n log n). So, theoretically, the awk method will eventually surpass the sort method because it's having to do less work (it's only checking for a previously seen key rather than sorting the entire file) but I'm not sure the crossover will be anywhere useful if the file doesn't contain many duplicates. Repeating it for a file containing lots of duplicates (same 128MB file size but contents are only the 7 letter words consisting of A or B, so only 128 possible entries):- $ time awk '!x[$0]++' y.out | wc -l 128 real 0m1.207s user 0m1.192s sys 0m0.016s $ time ( sort y.out | uniq ) | wc -l 128 real 0m14.320s user 0m31.414s sys 0m0.428s $ time ( sort -u y.out | uniq ) | wc -l 128 real 0m12.638s user 0m30.366s sys 0m0.188s Notice that "sort -u" doesn't do anything clever for files with lots of duplicates. So awk is much faster for files with lots of duplicates. No great surprises. When I get a chance I'll repeat it for a 1GB file and a 10GB file (with lots of duplicates otherwise the awk version will take far too long). 1. Example contents:- EPQKHPH DLJCROB WICVGQY MHWTPSR HMPNECN
- alexkus 14y agoThe awk run against a file containing almost no duplicates finished after over an hour (compared to 43sec for the sort method). $ time awk '!x[$0]++' x.out | wc -l 16759719 real 64m41.089s user 64m31.970s sys 0m3.136s Peak memory usage (given that it was a 128MB input file) was (pid, rss, vsz, comm): 8972 1239744 1246488 \_ awk So > 1GB for a 128MB input file.