3 ms·
My gut-ballpark was going to be around an order of magnitude, not two. Here's are two naive comparisons (granted, from the late aughts and not a direct comparis
by iheartmemcache 10y ago
My gut-ballpark was going to be around an order of magnitude, not two. Here's are two naive comparisons (granted, from the late aughts and not a direct comparison but a more general one) that show ~7-8x[0,1]. The overhead of PyStringObject is not trivial[2] (though the implementation details likely have changed between Py2.x and Python 3).
For things like building accumulators a set of data/log parsing rather than data munging (hits per hour or enumerative tasks), I'd imagine (g|n)?awk might hit your 100x since you'd just grab the fd and traverse being IO bound. I'm not sure how awk does it, but if it's just saving an accumulator value (or 10) in a register. Assuming x64-64 treats, say, an r3 fetch analogously to a fetch to ecx (err..rcx now I guess), rather than having to keep a full object in L1 cache, awk has a huge advantage.
---
N.b., if you're benching tasks like this, don't use 'time' and STDOUT and think you're getting real performance numbers. Your bottleneck (terminals can only render $x lines a minute, so the kernel call to write(STDOUT, ....) will be where you choke, not at the language. Also disk fragmentation would be another issue. Put the both the test file and the output file on a RAMdisk.) Cache flush with Something like sync; echo 3 > /proc/sys/vm/drop_caches (on Linux, I forget the BSD way of doing it) then `time benchmark.py /mnt/ramdisk1/file > /mnt/ramdisk2' over multiple runs, under various loads, with different data sets, etc
Another interesting thing to note is that comp arch is so advanced (I was reading a paper on formal verification of ISAs, and apparently even 12 dollar ARMs now have out-of-order instruction execution type stuff) that between the kernel scheduler and CPU optimizations, Python will likely benefit much more from disk-seek latency (effectively allowing PyStringObject allocation to occur while you're waiting for /dev/sd$n to return).
I'm certainly not an authority on rigorous benchmarks though - someone like Brendan Gregg please jump in!
[0] https://diamondinheritance.blogspot.com/2008/04/awk-vs-python-performance.html https://diamondinheritance.blogspot.com/2008/04/awk-vs-pytho...
[1] https://brenocon.com/blog/2009/09/dont-mawk-awk-the-fastest-and-most-elegant-big-data-munging-language/ https://brenocon.com/blog/2009/09/dont-mawk-awk-the-fastest-...
[2] http://www.laurentluce.com/posts/python-string-objects-implementation/ http://www.laurentluce.com/posts/python-string-objects-imple...
- doug1001 10y agonice one--i learned a few things, in fact (& caused me to realize once again, just how sloppy we are with benchmarking on my team)