3 ms·
On my system, I see the following: $ /usr/bin/time -f "%U\n" ./exp_volatile_a_b 0.13 $ /usr/bin/time -f "%U\n" ./exp_lapack 0.52 $ /usr/bin/time -f "
by twtw 8y ago
On my system, I see the following:
$ /usr/bin/time -f "%U\n" ./exp_volatile_a_b
0.13
$ /usr/bin/time -f "%U\n" ./exp_lapack
0.52
$ /usr/bin/time -f "%U\n" ./exp_openblas
0.32
I got LAPACK and BLAS via apt-get on Ubuntu 18.04, so whatever that means. I installed OpenBLAS from source, and it looks like the target detection decided to enable AVX2 (but not AVX512).
So the fully-unrolled version is still faster on the 5x5 matrix, which doesn't seem super surprising to me. I would expect a LAPACK implementation to have some overhead compared to a straight line solution to a problem this small.
- celrod 8y agoYes. This is easy to see in Julia, by comparing native arrays with StaticArrays.jl. Native arrays operations are linked to OpenBLAS, while many StaticArray operations are unrolled (will have to get to a computer to see if that includes 5x5 inversion). At the small sizes, BLAS is tens of times slower. Especially OpenBLAS (MKL does better at small sizes). There's a lot of overhead from things like determining optimal blocking structures.