4 ms·
Some more interesting points: - Switching to -O3 for gcc (instead of author's default -O2), the C version (exp_volatile_a_b) gets ~30% faster, bringing it down
by twtw 8y ago
Some more interesting points:
- Switching to -O3 for gcc (instead of author's default -O2), the C version (exp_volatile_a_b) gets ~30% faster, bringing it down to ~50% of the python-generated llvm - it appears to still be doing the full computation at runtime.
- Switching to -O3 for llc doesn't make any difference for the python-generated llvm version.
- With gcc 4.8 -O2 (instead of 7.3 -O2), I get ~0.06s - it looks like gcc 4.8 decides to inline everything, and gcc 7.3 doesn't.
- BeeOnRope 8y agoIt makes sense: gcc's -O3 is closer to LLVM/clang's -O2 in terms of big picture optimizations they enable such as vectorization and loop unrolling.