13 ms·
Removing duplicate lines from files keeping the original order with Awk
- founderling 7y agoAwk ' visited[$0]++' or how badly HNs automatic title formatter messes up awk commands :) Hint: You can edit it after you posted.
- laz_arus 7y agoDone, thanks for the hint :)
- founderling 7y agoYou are !done yet.
- deleted 7y ago[deleted]
- aidos 7y agoThis is a nice little run through of a real life example. Once you realise that awk's model is 'match pattern { do actions; }' everything makes a whole lot more sense.
- sohkamyung 7y agoAwk also supports BEGIN and END actions that take place at the start and at the end of execution. BEGIN might be used to initialise Awk variables or print initial messages, while END can be used to print a summary of actions at the end. See [1] https://www.grymoire.com/Unix/Awk.html#uh-1 https://www.grymoire.com/Unix/Awk.html#uh-1
- RBerenguel 7y agoYou can also use either BEGIN or END as the only entrypoints, essentially using AWK as you'd use any other programming language. Yes, sometimes this defeats the point of what AWK excels at, but it's good to know.
- mehrdadn 7y agoWhere I run into trouble with awk is gawk incompatibilities with the implementation on Mac. The gawk manual really sucks at telling you what exactly is an extension to the language, and I haven't been able to find a good source -- you just have to either guess and check, or cross-check against other ones' manuals (like BSD). Otherwise it's an amazing tool...
- asicsp 7y agoI thought the gawk book/documentation [1] did a good job of mentioning differences between various implementations, do you have an example? You might find this [2] helpful (oops, seems like it got deleted, see [3] - thanks @bionoid) [1] https://www.gnu.org/software/gawk/manual/gawk.html https://www.gnu.org/software/gawk/manual/gawk.html [2] https://www.reddit.com/r/awk/comments/4omosp/differences_between_gawk_nawk_mawk_and_posix_awk/ https://www.reddit.com/r/awk/comments/4omosp/differences_bet... [3] https://archive.is/btGky https://archive.is/btGky
- bionoid 7y ago[2] is available here: https://www.removeddit.com/r/awk/comments/4omosp/differences_between_gawk_nawk_mawk_and_posix_awk/ https://www.removeddit.com/r/awk/comments/4omosp/differences... Archive for posterity: http://archive.is/btGky http://archive.is/btGky
- mehrdadn 7y ago> do you have an example? Sure, try this: echo 1 2 | awk '{ print gensub(/1/, "3", "g", $1); }' The logical thing for them to do would be to mention in bold and/or big and/or red font under gensub's documentation that it's an extension (e.g. try nawk), whereas looking through it I don't see any mention at all: https://www.gnu.org/software/gawk/manual/html_node/String-Functions.html#String-Functions https://www.gnu.org/software/gawk/manual/html_node/String-Fu... If I may rant about this for a bit, GNU software manuals are generally rather awful (though they're neither alone in this nor is it impossible to find exceptions). They frequently make absolutely zero effort to display important information more prominently and unimportant information less so (if you're even lucky enough that they tell you the important information in the first place). Like if passing --food will accidentally blow up a nuke in your hometown, you can expect that if they documented it at all, they just casually buried it in the middle of some random paragraph. Their operating assumption seems to be that if you can't be bothered to spend the next 4 hours reading a novel before writing your one-liner then it's just obviously your fault for sucking so much.
- asicsp 7y agoI have a collection of such one-liners, for duplicates including how to form key for multiple fields, see [1] [1] https://github.com/learnbyexample/Command-line-text-processing/blob/master/gnu_awk.md#dealing-with-duplicates https://github.com/learnbyexample/Command-line-text-processi...
- laz_arus 7y agoThis repo is awesome, great work.
- seriousaccount 7y agoWow! Thanks for this! This is exactly what I was looking for :)
- asauce 7y agoWow this is great. I want to become more competent with text processing in the command line and this looks like a great place to start. Thanks for linking this repo!
- rusk 7y agoGot to love awk. My weapon of choice for ad hoc arbitrary text processing and data analysis. I’ve tried to replace it with more modern tools time and again but nothing else really comes close in that domain.
- RBerenguel 7y agoCompletely agree. Even Python, which has a very low barrier to entry to "read file, possibly csv, do something" has a barrier to entry. Column projections are one-liners in AWK, and aggregates and/or some stats can be a couple of lines in an AWK script proper. I've been replacing some ad-hoc bash scripts (nothing fancy, just a few if conditions and some formatting of outputs for a deployment) with some AWK, and it's so much handier to write (after 10 years I still can't remember if syntax) and read (it's a proper programming language) than bash edit: wrong markdown style
- vidarh 7y agoInteresting Ruby (MRI anyway) has command line options to make it act pretty similar to awk: -n adds an implicit "while gets ... end" loop. "-p" does the same but prints the contents of $_ at the end. "-e" lets you put an expression on the command line. "-F" specified the field separator like for awk. "-a" turns on auto-split mode when you use it with -n or -p, which basically adds an implicit "$F = $_.split to the while gets .. end loops. So "ruby -[p or n]a -F[some separator] -e ' [expression gets run once every loop]'" is good for tasks that are suitable for "awk-like" processing but where you may need access to other functionality than what awk provides..
- rusk 7y agoThanks, I actually looked at trying to make ruby do my awk work a few years ago. I'll take a look again.
- asicsp 7y agoI'd say it is more similar to perl than awk for options like -F -l -a -n -e -0 etc. And perl borrowed stuff from sed, awk, etc I have a collection for ruby one-liners too [1] [1] https://github.com/learnbyexample/Command-line-text-processing/blob/master/ruby_one_liners.md https://github.com/learnbyexample/Command-line-text-processi...
- omh 7y agoAwk is wonderful. It's an odd way to write programs, but for quick one-off processing tasks it almost can't be beaten. Somewhat related blog post which I like to refer people to: "Command-line Tools can be 235x Faster than your Hadoop Cluster" https://adamdrake.com/command-line-tools-can-be-235x-faster-than-your-hadoop-cluster.html https://adamdrake.com/command-line-tools-can-be-235x-faster-...
- chasil 7y agoWhy do you find it odd? I find it to be the very best introduction to C-style control structures. Chapter 2 of The AWK Programming Language has incredible benefits for a novice. https://archive.org/download/pdfy-MgN0H1joIoDVoIC7/The_AWK_Programming_Language.pdf https://archive.org/download/pdfy-MgN0H1joIoDVoIC7/The_AWK_P...
- jolmg 7y agoThey probably refer to how it has the top-level as <line-condition> { <code> } That's unusual enough among programming languages to call it odd. Being able to do stuff like if (/some-pattern/) { ... and have the regex be evaluated like a condition where it matches with the current line implicitly is also pretty unique.
- funkymike 7y agoif (/some-pattern/) { ... This isn't really unique when you consider perl. while (<>) { if (/pattern/) { This does the same. Awk simply has the implicit loop.
- jolmg 7y agoYes, I think Perl based that on Awk, but then those are the only 2 languages I know that support something like that. That's still very unique. On the implicit loop, along with Ruby, they're the only 3 languages I know that support something like that. That's also pretty unique, and Awk is the only one that has the implicit loop as a requirement.
- augustk 7y agoAnd here is the ungolfed version: awk '{ if (! visited[$0]) { print $0; visited[$0] = 1 } }'
- zufallsheld 7y agoMuch more readable and understandable.
- RBerenguel 7y agoYou can always write AWK in a file and read the script with -f, making it fully readable (and AWK is quite a pleasantly readable and surprisingly versatile language to write at that point)
- asicsp 7y agoto add to this, if you've coded one-liner first, you can convert to script using -o option for ex: awk -o '{ORS = NR%2 ? " " : RS} 1' gives (default output file is awkprof.out) { ORS = (NR % 2 ? " " : RS) } 1 { print $0 }
- nerdponx 7y agoThat I did not know, great tip, thank you!
- RBerenguel 7y agoI wasn't aware of this, might come handy for "one liner edge cases".
- augustk 7y agoOr maybe better: awk '! visited[$0] { print $0; visited[$0] = 1 }'
- asicsp 7y agoyet another version: awk '!($0 in seen); {seen[$0]}'
- hk__2 7y agoThe tradeoff of this solution is it stores all (unique) lines in memory. If you have a large file with a lot of unique lines you might prefer using `sort -u`, although it doesn’t keep the order.
- omh 7y agoDoesn't sort require keeping lines in memory as well? In fact, doesn't sort keep all lines in memory, whereas this awk solution just keeps the unique lines.
- ot 7y agoNo, most implementations of sort use a bounded memory buffer and an external-memory algorithm, spilling to disk.
- chii 7y agoI don't believe that - how does it spill to disk? TMP directory? You can still sort if you don't have any disk space (until you run out of swap?) from what I recall.
- chasil 7y agoThe man page for sort has parameters to adjust the location of the temporary files. These are only used when the allowed memory buffer is exhausted.
- JdeBP 7y agoYour recollection is either limited in what implementations you have encountered, or faulty. * https://unix.stackexchange.com/a/450900/5132 https://unix.stackexchange.com/a/450900/5132
- ot 7y agoThis is mentioned in the article, together with a method to keep the order by decorating with the line number before sorting.
- lifeisstillgood 7y agoThere was a twitter thread ages ago where someone had written a collection of (php?) utilities - and the twitterer posted a laughing slap down saying why write a utility when this one liner and that one liner will do? There was a lot of push back - and this article is a good example of why If I wanted to remove duplicate lines from a file I would almost certainly not use awk I have never spent the time to get good enough with the whole new and different languge of awk, and am unlikely to need to (my large scale file processing needs seem small, and if I do it's almost always in context of other processing chains - so a normal languge like python would be the natural choice I could whip up something like this in python in a less time than it would take to google the answer, read up why the syntax works that way and verify I have not mistyped anything on a few test files. Basically using awk takes me out of my comfort zone - for a one off task it loses me time, for a production like repeat task I am going to reach for a slew of other solutions. I mean the title of this page loses the exclamation mark - and it took me two goes to spot it.
- asicsp 7y agoI don't get your point, it seems like you do not often use cli text processing tools Just like Python, there are users who use cli and are comfortable using grep/sed/awk/sort/etc
- nerdponx 7y agoThe point is that the people who do use such tools tend to have a derisive attitude towards those who don't, and that the derisiveness is completely unwarranted.
- reacweb 7y agoI have taken as rule to use awk only for trivial tasks and to switch to perl as soon as the syntax is slightly beyond my usual use cases. In perl, I would do: perl -nle 'print unless exists $h{$_};$h{$_}++' < your_file
- asicsp 7y agoperl borrowed stuff from awk, so you could also do perl -ne 'print if !$seen{$_}++'
- rlonstein 7y agoPlay a round of perl golf? perl -pne '$_=$#$_++?$_:""' I'm rusty at this but shaved off six chars, five if you count the 'p' added to switches.
- asicsp 7y agoedit: this is failing if a line is repeated more than once, see @showdead's excellent explanation the shortest I've got so far is perl -lnE'say if!++$#$_' --- I don't understand what's happening with $#$_ but seems like something I should look into, thanks :) you could remove n as p is used and would be same no. of characters as perl -ne 'print if $#$_++' you could save one more by removing space between e switch and single quote
- reacweb 7y agoI have never seen this $#$var trick and google is not the friend of perl operators. Do you have any explanation ? perl -ne 'print if!++$#$_' seems to work also
- showdead 7y agoIf you have array @foo in perl, $#foo is the index of the last element of the array, which is just the size of the array -1. So if @foo is undefined, $#foo is -1. Using a variable instead of 'foo' is a symbolic reference, so this is effectively using the symbol table as the associative array. This means that this solution also gets it wrong if your file contains a line that matches the name of a built-in variable in perl. That would be tough to debug! If your file contains This is the first line of the file then during execution of ++$#$_ the result is the same as if you had written ++$#{This is the first line of the file} So the variable @{This is the first line of the file} goes from undefined to an array of length 1, turning $#{This is the first line of the file} to 0. Incidentally, this is why the snippet fails to work for a line repeated more than once: for each occurrence of the expression, the value returned is in the sequence -1, 0, 1, 2, 3, ... so it is only false for the second occurrence. Using preincrement instead of postincrement means the values returned are 0, 1, 2, 3, ... which means that inverting the test makes it false for every occurrence after the first.
- triangleman 7y agoSo then, what is the one liner to preserve the filename rather than get a new deduped.txt? Also how do you apply that command to the next file using shell history?
- gcmeplz 7y agoUse sponge! https://linux.die.net/man/1/sponge https://linux.die.net/man/1/sponge awk '!n[$0]++' fileName | sponge fileName
- Vogtinator 7y agoOr using gawk, gawk -i inplace 'foo {bar}' file
- dredmorbius 7y agodedupe.awk <file >file.tmp && mv file.tmp file Multifile versions vary, I'd prefer listing them out, alternatively you could read from a command output (ls, find, etc.) with a 'while read; do ... done' loop: for f in file1 file2 file3 do dedupe.awk <$f > ${f}.tmp && mv ${f}.tmp $f done If you want to apply to specific files on an ad hoc basis, you could wrap the whole thing in a shell function with filename or list as a parameter. Or 'gawk -i' as suggested. Properly using tempfile would also be an improvement.
- julienfr112 7y agoI was wondering how awk work internally. Does it compile the script then run it ? Is it bytecode or lower level ? A finite state machine ?
- CodeArtisan 7y agoWith GNU awk, scripts are compiled to bytecode then interpreted with a big switch/case loop. edit: https://git.savannah.gnu.org/cgit/gawk.git/tree/interpret.h https://git.savannah.gnu.org/cgit/gawk.git/tree/interpret.h
- chasil 7y agoThe mawk version of the language will output C, if I remember correctly. It is the fastest AWK, and supports some of the GNU extensions. "The One True AWK" from Brian Kernighan (that is still the system AWK in OpenBSD) switched from a yacc implementation to a custom parser sometime within the last decade (fairly recently). Busybox also has an awk; I'm not sure what they do. GNU awk is elsewhere reported to be interpreted bytecode.
- snaky 7y agoDennis Ritchie about yacc history details > In some ways the interesting thing is that the parser (probably for B, couldn't have been C based on radiocarbon dating evidence) was tiny and dead simple using recursive descent for most parts, a precedence table for expressions. But out of the intellectual culture-meets-culture encounter, an enduring tool was created. https://yarchive.net/comp/handwritten_parse_tables.html https://yarchive.net/comp/handwritten_parse_tables.html
- fanf2 7y agoLooks like OpenBSD awk still uses yacc, like other copies of bwk‘s one true awk https://cvsweb.openbsd.org/cgi-bin/cvsweb/src/usr.bin/awk/#dirlist https://cvsweb.openbsd.org/cgi-bin/cvsweb/src/usr.bin/awk/#d...
- benhoyt 7y agoThey parse the script to a parse tree (abstract syntax tree) and then either interpret that directly (tree-walking interpreter) or compile to bytecode and then execute that. The original awk ("one true awk") uses a simple tree-walking interpreter, as does my own GoAWK implementation. gawk and mawk are slightly faster and compile to bytecode first. If you're interested, you can read more about how GoAWK works and performs here: https://benhoyt.com/writings/goawk/ https://benhoyt.com/writings/goawk/
- deleted 7y ago[deleted]
- bigato 7y agoThis is very memory intensive. Which may not matter if the data volume is small enough. But it is also a bit hard to understand, at least not so obvious at first sight. For most use cases sort -u would be ideal and way simpler to understand, if you don't mind having an ordered file at output.
- beefsack 7y ago`sort` would also be memory intensive would it not?
- chasil 7y agoNo, sort will use intermediate temporary files instead of exhausting your ram.
- mitnk 7y ago> This is very memory intensive. Only for ones not familiar with awk. It would make a lot of sense after you understand how awk works (as the article explains).
- anc84 7y agoThat does not make any sense. If it is memory intensive depends on awk, not on the person being familiar with it. Sois it memory intensive or not?
- chasil 7y agoThe example AWK script will build an array of every unique line of text in the file. If the file is large, and mostly unique, then assume that a substantial portion of the file will be loaded into memory. If this is larger than the amount of ram, then portions of the active array will be paged to the swap space, then will thrash the drive as each new line is read forcing a complete rescan of the array. This is very handy for files that fit in available ram (and zram may help greatly), but it does not scale.
- mitnk 7y agoThe man page of (n)awk [0][1] is surprisingly short and readable. [0] `man awk` on mac [1] online version https://www.mankier.com/1/nawk https://www.mankier.com/1/nawk [2] gawk's man page works great as a reference https://www.mankier.com/1/gawk https://www.mankier.com/1/gawk
- hjk05 7y agoOn a previous post people were complaining that math wasn’t as clear as code. I’d argue that this this is exactly the kind of code-like clarity math notation provides you. I makes perfect sense, but only after 2 full pages describing what’s going on in one line.
- deleted 7y ago[deleted]
- oftenwrong 7y agoSee also, 'nauniq' (non-adjacent uniq), an implementation of this text-processing task as a full utility, with some convenient options for reducing memory usage: https://metacpan.org/pod/distribution/App-nauniq/script/nauniq https://metacpan.org/pod/distribution/App-nauniq/script/naun...
- deleted 7y ago[deleted]
- kazinator 7y agoYou need "gawk -M" for this for bignum support, so visited[$0]++ doesn't wrap back to zero, otherwise it is not correct for huge files with huge numbers of duplicates. The portable one-liner that doesn't suffer from integer wraparound is actually awk '!($0 in seen) { seen[$0]; print }' which can be golfed a bit: awk '!($0 in s); s[$0]' $0 in s tests whether the line exists in the s[] assoc array. We negate that, so we print if it doesn't exist. Then we unconditionally execute s[$0]. This has an undefined value that behaves like Boolean false. In awk if we mention an array location, it materializes, so this has the effect that "$0 in s" is now true, though s[$0] continues to have an undefined value.
- zimpenfish 7y ago> huge files with huge numbers of duplicates At least on the stock MacOS awk, you can get up to 2^53 before arithmetic breaks (doesn't wrap, just doesn't go up any more which means the one-liner still works.) > echo '2^53-1' | bc 9007199254740991 > seq 1 10 | awk 'BEGIN{a[123]=9007199254740991;b=a[123]}{a[123]++}END{print a[123],b,a[123]-b}' 9007199254740992 9007199254740991 1 Even with one character per line, you'd need an 18PB file before you got to this limit, afaict.
- deleted 7y ago[deleted]
- jancsika 7y agoMe: Wow, that associative array looks very powerful. Is there a way to leverage it to do something useful like convert curl-obtained JSON array of file patches from Github's API to the mbox format that `git am` expects? Unix: No, that JSON data is too structured. But if you have a more error-prone format like CSV I can show you a neat trick to filter your bowlers by number of spares.
- james_s_tayler 7y agocan't u just process JSON data with jq? https://stedolan.github.io/jq/ https://stedolan.github.io/jq/
- jancsika 7y agoI'd still need to find a way to massage the data into the mbox format because that is the ancient format that git understands. I'm not saying that there isn't a way to do that. Only that it can only be done poorly with a big ugly (and probably buggy) spaghetti script that looks nothing like what the expressive demo suggests it should look like.
- james_s_tayler 7y agoIf you're getting a bunch of stuff with curl from GitHub can't you just use curl to get the patches directly from github? Append .patch to the end of a pr or commit and it spits out the mbox formatted patch. https://github.com/jiphex/mbox/commit/f139c575e306a1691a31d8c2f4b44f48984b1267.patch https://github.com/jiphex/mbox/commit/f139c575e306a1691a31d8...
- jancsika 7y agoIf you do that with curl it will redirect to the login page. Github obviously wants me to use their API. I assume this wasn't always the case as the use case I'm referencing is a build script I'm debugging.
- 7y ago
- pvaldes 7y agoopen the file with emacs menu edit -> select all press 'escape' and 'x' keys together and write: delete-duplicate-lines done
- aabbcc1241 7y agoIf I know this, I wouldn't make https://github.com/beenotung/uniqcp https://github.com/beenotung/uniqcp