3 ms·
Has anyone reproduced these results? FWIW, I can't. $ cd wordsandbuttons/exp/python_to_llvm/exp_c $ make $ /usr/bin/time -f "%U\n" ./exp_volatile_a_b
by twtw 8y ago
Has anyone reproduced these results?
FWIW, I can't.
$ cd wordsandbuttons/exp/python_to_llvm/exp_c
$ make
$ /usr/bin/time -f "%U\n" ./exp_volatile_a_b
0.13
$ cd wordsandbuttons/exp/python_to_llvm/exp_embed_on_call
$ make
$ /usr/bin/time -f "%U\n" ./benchmark
0.16
Intel(R) Core(TM) i7-7800X CPU @ 3.50GHz, gcc (Ubuntu 7.3.0-27ubuntu1~18.04), clang version 6.0.0-1ubuntu2, LLVM version 6.0.0
- twtw 8y agoSome more interesting points: - Switching to -O3 for gcc (instead of author's default -O2), the C version (exp_volatile_a_b) gets ~30% faster, bringing it down to ~50% of the python-generated llvm - it appears to still be doing the full computation at runtime. - Switching to -O3 for llc doesn't make any difference for the python-generated llvm version. - With gcc 4.8 -O2 (instead of 7.3 -O2), I get ~0.06s - it looks like gcc 4.8 decides to inline everything, and gcc 7.3 doesn't.
- BeeOnRope 8y agoIt makes sense: gcc's -O3 is closer to LLVM/clang's -O2 in terms of big picture optimizations they enable such as vectorization and loop unrolling.