9 ms·
So, I don't use awk right now. This answers "how to awk" more than "why use awk" for me (understandable given the 20 min claim). Does anyone have concrete, pra
by bcyn 6y ago
So, I don't use awk right now. This answers "how to awk" more than "why use awk" for me (understandable given the 20 min claim).
Does anyone have concrete, practical examples of use cases where awk made your life much easier?
- ggm 6y agoAwk '{ print $2 }' does LWSP gobbling.. cut -d' ' -f2 doesn't and this difference alone makes awk useful to me on a daily basis. Awk has a hash like perl which is very efficient. I use an awk expression to print uniq as they come in counted through the hash insert on new instead of uniq which prints at end. Awk count unique over 300,000,000 ips was as fast as perl and python and smaller memory footprint
- swimfar 6y agoLWSP=Linear White Space "Linear white space is: any number of spaces or horizontal tabs, and also newline (CRLF) if it is followed by at least one space or horizontal tab." [1] I don't know what LWSP gobbling is, though. Google didn't help. [1] https://stackoverflow.com/questions/21072713/what-exactly-is-the-linear-whitespace-lws-lwsp https://stackoverflow.com/questions/21072713/what-exactly-is...
- pstuart 6y agoIt means a contiguous collection of arbitrary white space characters are considered to be a single delimiter. With the parent example it means that column #2 could have any of number whitespaces between it and column #1.
- ggm 6y agoyes. whereas cut, considers ' ' (two spaces) as an instance of a blank column separated by the -d' ' space instance. Its like ,, in CSV being a blank column.
- skrebbel 6y ago> Awk '{ print $2 }' does LWSP gobbling That's like answering "why use Haskell?" with "Because its monads are monoids in the category of endofunctors"
- ggm 6y agoThis made me very happy, since it goes to the 'you can either use awk or explain awk but not both' problem which is definitional in monads in Haskell. I thought LWSP was a thing. I forgot not everything is a thing. And gobbling is maybe very jargon ish. Python sys.stdin.rstrip().split() comes to mind.
- chme 6y agoIf you just use awk for `{ print $2 }` then I would still prefer `tr -s ' ' | cut -d ' ' -f2`, since both are part of coreutils and with `awk` you would add a additional dependency to your script.
- asicsp 6y agothat is not equivalent because awk will remove trailing/leading space/tab/newlines (newlines come into play with a different record separator) whereas, tr will still leave a trailing/leading space for example: $ echo ' a b c ' | tr -s ' ' | cut -d ' ' -f2 a $ echo ' a b c ' | awk '{print $2}' b and by default awk splits on space/tab/newlines, whereas in cut example above, you get only space as delimiter cut has its uses and will be faster than awk, but it depends on the problem being solved
- chme 6y agoYou are right it depends on the specific use case. It possible to use `xargs -L 1` to trim the separators, but then you would also add `findutils` as deps. I just wanted to point out that keeping the dependencies of scripts in mind when programming is also important.
- em500 6y agoWhy would being part of coreutils matter? Awk is part of POSIX, just like tr and cut. Even in very constraint environments you can count on it via busybox.
- chme 6y agoA lot of stuff is part in POSIX, but not all of that is available on every system. Also busybox configurations can vary a lot.
- ggm 6y agoA perpetual argument is use of too many pipe separated distinct commands. The argument suggests it's lazy to use sed | awk | grep type pipe runs because in all probability sed or awk alone could have done it and you incurred two excess fork/exec() and therefore consumed kernel and userspace beyond your need. The usual rejoinder is "get nicked"
- arexxbifs 6y agotr will handle the whitespace for you and is generally faster[0] (but awk is faster to type and reads better). [0] https://datagubbe.se/cutvawk/ https://datagubbe.se/cutvawk/
- xorcist 6y agoFor the simplest uses, cut is more readable. If argument expansion is required, consider using read. That way the argument is given a name. Use "read a b c < x" instead of "b=$(cat x | awk '{ print $2 }')". Just remember to pipe to read with care, as the right hand part of a pipe is a subshell in which variables are local. So "read a b c < <(echo 1 2 3)" works, "echo 1 2 3 | read a b c" doesn't. Another way to expand arguments is to simply define a shell function and use $2. Something like "process_line() { echo $2 }". When you have to reach for something like awk, your script would probably improve by being mostly awk. In which case most people are probably looking at perl or python anyway. Even if awk is a nice language, the arrival of perl mostly killed it. It is not wrong to say perl was the next version of awk.
- na85 6y agoI contribute to a niche game server project and used awk to generate c# class files from an org mode document that I wrote to specify them.
- bcyn 6y agoWhat about awk makes that easier than, say, a Python script?
- na85 6y agoProcessing text files line by line is trivial in awk. Maybe it's easy in Python too but at least in awk you don't have to worry about the 2 vs 3 silliness. Plus I think semantic indentation is not sane language design. Plus you can pipe it to another command with ease.
- 7thaccount 6y agoAwk automatically operates on every line of a text file, so you can easily use it as part of a pipeline in Linux. I'll usually use grep, cut, sort, and awk all together with one line of code and no need to open an editor. Python takes a lot more code to do these basic things. However, Python is much better once the complexity grows past a certain point.
- asicsp 6y agodepends upon the usecase, for relatively simple column filtering, substitution, etc, awk script would be briefer than python equivalent and most likely be faster as well if it becomes lengthy (and again depends on features required), then Python is likely to have inbuilt/3rd-party libraries to make it easier to write and maintain if it is a question of constructing command line one-liners and using it as a part of other cli tools, then awk wins easily for example: awk -F'\\W+' -v OFS=, '{print $NF, $2}' input.txt prints last column and second column, where non-word characters form the field separator and comma is used as output field separator
- gen220 6y agoI use awk to write csv-parsers for my personal finance setup. They take transaction logs and convert them into a ledger format (https://www.ledger-cli.org/ https://www.ledger-cli.org/). There's some pretty trivial if/else branches based on regexp's, which awk is pretty good at expressing. I'd normally write this sort of thing in python. The awk program is usually smaller than a comparable python program, but for me the main selling point is that awk programs a more "UNIX-pipe-native" than python programs are. At some point, I'll probably rewrite these programs in a more "serious" programming language, once the file sizes get too big or something, but for now they're working great and are easy to extend.
- mauvehaus 6y agoI've used it for this too, but things tend to go south pretty quick if you have escaped or quoted commas. That, unfortunately, is where I usually break down and pull out python.
- gen220 6y agohehe, I hit that bump in the road too, but maybe I was too stubborn. FWIW, you can get around it. I just took a peek because I couldn't remember how to do it off the top of my head. It looks like I had to use gawk -F ',' -v FPAT='([^,]+)|("[^"/]+")' to get the behavior you're looking for. Seems like I nabbed it from https://www.gnu.org/software/gawk/manual/html_node/Splitting-By-Content.html https://www.gnu.org/software/gawk/manual/html_node/Splitting... Agreed that this is the sane point at which to pull out python.
- mauvehaus 6y agoI hadn't run across that one yet. Thank you! The risk of that seems to be if CSV allows embedded escaped quotes in a quoted string. Does it? I don't know. And CSV is pretty loosely defined. For most people it's probably "whatever Excel emits or ingests". And I think that's why we're on the same page about pulling out python and using a module where somebody has explored what the corner cases are and dealt with them for us already.
- asicsp 6y agoHere's some articles * https://blog.jpalardy.com/posts/why-learn-awk/ https://blog.jpalardy.com/posts/why-learn-awk/ (also discussed on HN: https://news.ycombinator.com/item?id=22108680 https://news.ycombinator.com/item?id=22108680) * https://adamdrake.com/command-line-tools-can-be-235x-faster-than-your-hadoop-cluster.html https://adamdrake.com/command-line-tools-can-be-235x-faster-...
- AceJohnny2 6y agoI use awk when I need something a bit more powerful than "grep"/"cut", but don't want to pull out the big guns with Perl. For example, recently I needed to print out certain fields of an output, but only for a given subset. So I used Awk to create a simple state machine (enable when I see the start of the subset, disable at the end), and print the fields of interest.
- AceJohnny2 6y agoOr you could use it to solve the Towers of Hanoi game ;) https://rosettacode.org/wiki/Towers_of_Hanoi#AWK https://rosettacode.org/wiki/Towers_of_Hanoi#AWK
- rmetzler 6y agoThe grep/cut replacement is exactly what I use awk the most for. I use it for other things too but this stands out. awk just fixes the problem when you need to have more than one field from the input and maybe a different string between them. I guess this can be done with cut too, but for me it's simpler this way. Also I use awk often in scripts together with fzf to interactively switch contexts e.g. in cloud provider CLIs.
- billjings 6y agoI wrote my answer to this a few years ago: https://www.bignerdranch.com/blog/a-crash-course-in-awk/ https://www.bignerdranch.com/blog/a-crash-course-in-awk/ A tl;dr for the following tl;dr: it's great for quick-and-dirty processing of logs. The tl;dr is that AWK is simple language that lets you use a line-oriented event driven programming model to process text. This abstraction is simple enough to be readable and maintainable, but powerful enough to parse and process text with line-oriented structure. The example linked elsewhere is worth studying, because it truly is one of the prettiest pieces of programming I've seen in a long while: https://c2.com/doc/expense/ https://c2.com/doc/expense/ If you don't know anything about AWK, here's all you need to know to grok this: 1. An AWK program is a list of [event] { code } pairs. All pairs are run in sequence on each line of input; if the "event" evaluates to true or is a matching regex, the matching code runs. 2. { code } on its own runs unconditionally. 3. The input is automatically whitespace delimited into fields; $1 refers to field 1, $2 to field 2, etc 4. NF is a special value that yields the number of fields 5. Assigning to fields replaces the text in those fields. 6. ($1+0) != 0 implicitly converts $1 to an integer; if it fails, the value is 0. There is a lot of implicit loose typing in the language, and usage of undefined variables is idiomatic. Functions aren't easy to use, either. So it's not well-suited for programming in the large, or even the medium. But for the scale of programming seen in that link, it's truly a wonderful and simple power tool.