5 ms·
I use sed & awk all the time. It is an invaluable tool while debugging issues, extracting fields from log files etc. I am not dissing python or perl, I use pyth
by 0xFFFE 9y ago
I use sed & awk all the time. It is an invaluable tool while debugging issues, extracting fields from log files etc. I am not dissing python or perl, I use python extensively as well. But while you are in the middle of an incident, hard to beat a quick oneliner, like this random example.
awk -F":" 'BEGIN{total=0}{if($3>240)total+=$3}END{print total}' /etc/passwd
- indescions_2017 9y ago>>> I use sed & awk all the time. Me too. For example, number of unique IP requests in a log file in milliseconds ;) $ awk '{ print $1 } ' caddy_log | sort | uniq | wc -l But rarely compose my own from scratch. It's mostly copy paste. And store in admin bin for future use.
- dietr1ch 9y agoThere's `sort -u` which should be able to avoid some work.
- ninkendo 9y agoDumb side note: I always end up using `awk '{print $1}'` instead of `cut -f1` because cut's field separation defaults are unwieldy in that it doesn't intelligently consider any whitespace to be a field separator, which is what I need 99% of the time.
- ptman 9y agosometimes (if there are not mixed spaces and tabs) tr -s ' ' < input | cut -f1 helps
- asicsp 9y agoNo need so many commands :) $ cat duplicates.txt abc 7 4 food toy **** abc 7 4 test toy 123 good toy **** $ awk '!seen[$2]++' duplicates.txt abc 7 4 food toy **** $ awk '!seen[$2]++{cnt++} END{print +cnt}' duplicates.txt 2
- dredmorbius 9y agoIf you just want to count unique IPs, you can get that in awk alone and avoid sort: awk '{ip[$1]=1}; END{ print length(ip) }' Use the hash tables directly, assuming you've got the memory. That should be slightly faster as it avoids the sort.
- asicsp 9y agonice one also, {ip[$1]} is enough...
- dredmorbius 9y agoThanks, I was trying to think of the fastest possible assignment.
- c4n4rd 9y agoNo need to spawn another process(uniq), you can use sort -u ;-)
- kbenson 9y agoFor me, the benefit of a trailing uniq instead of a sort -u is that I'm invariably building towards a "sort | uniq -c | sort -n" anyways.
- ajross 9y agoInterestingly, perl was designed to fill exactly that role, with a syntax that was at the time considered cleaner and more intuitive than awk's (and of course with a regex backend that was much more powerful). In this case, something like: perl -e 'while(<>) { $tot += (split)[2]; }; print "$tot\n"' It's funny to see that in perl's decline, the stuff it did well is being forgotten and resurrected via tools it at one point had mostly replaced.
- asicsp 9y agoPerl does have command-line options to do that :) perl -F: -lane '$total+=$F[2] if $F[2]>240; END{print $total}' /etc/passwd
- ajross 9y agoOf course it does, the implicit looping and BEGIN/END blocks (I'll concede I had actually forgotten about the implicit splitting) were in fact deliberately designed to emulate awk. Nonetheless I don't think that really changes the point much. The core language syntax is simple enough that perl one-liners can be written without that and still get the point across.
- majewsky 9y agoLet me golf this. :) perl -F: -anE '$tot+=$F[2]}{say$tot'
- jcims 9y agoThis made me miss perl (but ya forgot $F[2]>240)
- majewsky 9y agoAh, I was going by the snippet in the parent, not the original one in the grandparent. That here should do it: perl -F: -anE '$_=$F[2];$t+=$_if$_>240}{say$t' I wanted to be clever by doing `$t+=$_*($_>240)`, but that's actually one byte larger. :(
- asicsp 9y agoyou could simplify by taking advantage of default value which is `0` in numeric context awk -F: '$3>240{total+=$3} END{print +total}' /etc/passwd
- vram22 9y agoThe fragment BEGIN{total=0} can be skipped (at least in the awk's I've used, not sure if some more recent / strict awk differs), because awk initializes the variable total to 0. I remember this because I do this all the time to quickly get the (non-recursive) sum of the sizes of the files in a directory - it's pretty much muscle memory from a while now: ls -l | awk '{ s += $5 } END { print s/1024 " KB" }' For recursive size, one can use ls -lR or the du command with various options according to need.
- kazinator 9y agoBut that only works because in the final print, you have done arithmetic with s. So in other words, we can take out the BEGIN block, but then we must remember to change print total to print total + 0. Also, using uninitialized variables is basically a code golfing stupidity that will bite you in any halfway complicated program. GNU Awk has a useful --lint argument which spots uses of uninitialized variables. If you make it habit to write code that way, if you then use --lint for finding a bug, you have to deal with false positives.
- vram22 9y ago>Also, using uninitialized variables is basically a code golfing stupidity that will bite you in any halfway complicated program. Nonsense. Not if you know what you are doing, and used it in a known way, which is what I did. The code I wrote works. I tested it on Linux before posting it. Also, such a usage (skipping the initializer) is mentioned (IIRC) in the classic Kernighan & Pike book "The Unix Programming Environment" (still a great resource, though not updated for modern Unix/Linux features), which is where I learned it from, years ago (and hence why I qualified my statement by saying it may not work in more strict or modern awk versions). Fine to talk about other variations but it does not mean that my variation is wrong. Don't try to read my mind. My intention was not code golfing. Was just sharing some fun info. It's not a big deal to keep the initializer either, I'm quite aware of that.
- kazinator 9y agoYou do not know 100% what you're doing. Evidence being, in the grandparent comment you wrote "[...] because awk initializes the variable total to 0" which isn't how awk works at all. Your intention can be understood as the promotion of code golfing, as evidenced by these words: The fragment BEGIN{total=0} can be skipped by which you're clearly encouraging that other coder to make their code shorter by removing an initialization that works fine.
- feelin_googley 9y ago#!/bin/sh x=0;while true;do read a; test ${#a} -gt 0||exec echo $x a=${a#*:\*:};n=${a%%:*}; test $n -le 240||x=$((x+n)); done < /etc/passwd
- feelin_googley 9y ago#!/bin/sh IFS=:;x=0;while true;do read a b c d; test ${#a} -gt 0||exec echo $x test $c -lt 240||x=$((x+c)); done < /etc/passwd