3 ms·
Maybe I wasn't clear about this distinction. Python is an interpreted language. Even if you are matching a "compiled" regex, you are repeatedly interpreting the
by rspeer 8y ago
Maybe I wasn't clear about this distinction. Python is an interpreted language. Even if you are matching a "compiled" regex, you are repeatedly interpreting the line that says to match it.
Awk is a compiled language. Your Awk script is compiled once and applied to every line of your file at C-like speeds. It is way faster than Python.
If you learn to use Awk well, you will start doing things with data that you wouldn't have had the patience to do in an interpreted language.
- sametmax 8y agoTechnically, python is compiled to pyc and not reparsed again and again. What'd more, string operationd are all implemented in c. Unless you use pypy, which is faster.
- rspeer 8y agoTrust me. Awk is much faster. I say this with no disrespect for Python, it's just you should know the tradeoffs of the tools you use. It does not matter if the individual operations in Python are implemented in C. It will still be much slower at looping over lines of a file than a compiled language designed decades ago for that exact purpose.
- sametmax 8y agoI'm really unsure about my benchmarking capabilities, yet I wanted to have a general idea of the order of magnitude we are talking about. Reading your comment, I was expecting at least a X5 speed up by using awk. I have a lot of Python file in one dir: $ find . -iname "*.py" | wc -l 10429 Finding them and cating them all takes about 0.3 secs: $ time find . -iname "*.py" | xargs cat {} > /tmp/cat.out ... real 0m0.344s user 0m0.140s sys 0m0.175s So to have something simple that takes a bit of time, I tried to get all lines starting with "print", and output the first thing after that. I'm really bad at awk, so I don't know if there is a better way. I went for the most obvious thing for me: $ time find . -iname "*.py" | xargs cat {} | awk '/^print/{print($2)}' > /tmp/awk.out ... real 0m1.111s user 0m1.165s sys 0m0.368s Now, with Python, it's definitely not as easy to type. You have to get a script like. awk wins the expressivity metrics for this use case: import sys for x in sys.stdin: if x.startswith('print'): try: print(x.split()[1]) except IndexError: print('') # to match awk behavior But as for performance, I don't get the huge boost in perfs you are talking about: $ time find . -iname "*.py" | xargs cat {} | python /tmp/test.py > /tmp/python.out ... real 0m0.762s user 0m0.862s sys 0m0.347s I do get the same output though: $ cmp /tmp/python.out /tmp/awk.out && echo "yes" yes
- rspeer 8y agoThanks for running these statistics. Your experience is certainly different from mine. Is your Python actually PyPy? Here's what I observed: I had a simple text-processing tool I needed that I call "countmerge", which just merges adjacent lines with the same key and adds up their corresponding values. I needed to run it on a lot of large files. I first wrote it in Python, where it was a significant bottleneck compared to the steps that came before it (split, sort, uniq -c). Eventually I rewrote it in Rust [1], and it was at least 5 times faster, at the expense of a fair amount more low-level code. But then rewriting it in awk [2] turned out to be as fast as what I wrote in Rust, possibly inconclusively faster. [1] https://github.com/rspeer/countmerge https://github.com/rspeer/countmerge [2] https://gist.github.com/rspeer/60c87dca1ab550326f8bd6d086452613 https://gist.github.com/rspeer/60c87dca1ab550326f8bd6d086452...
- sametmax 8y agoOk, I generated a file similar to your test case >>> with open('data.txt', 'w') as f: ... for l in string.ascii_uppercase: ... for x in range(0, random.randint(1, 100000)): ... f.write('Key {}\t{}\n'.format(l, random.randint(0, 100))) With this script: import sys old_key = total = 0 for line in sys.stdin: key, value = line.split('\t') if old_key != key: old_key = key total = 0 print(key, value, end="") total += int(value) I get: $ <data.txt time python3 test.py Key A 2 Key B 87 Key C 58 Key D 64 Key E 29 Key F 25 Key G 2 Key H 74 Key I 17 Key J 37 Key K 97 Key L 77 Key M 19 Key N 74 Key O 33 Key P 61 Key Q 67 Key R 23 Key S 4 Key T 70 Key U 25 Key V 15 Key W 35 Key X 17 Key Y 31 Key Z 18 1.03user 0.01system 0:01.05elapsed 99%CPU (0avgtext+0avgdata 9564maxresident)k 0inputs+0outputs (0major+1100minor)pagefaults 0swaps But I can't manage to get the awk version working. It only prints one line on Ubuntu 16.04: $ <data.txt awk -f ./countmerge.awk Key 0 So I can't check it.