3 ms·
This kind of matches the performance I recorded when numpy is linked to vecLib for large matrix matrix multiplication in float32: https://gist.github.com/ogris
by ogrisel 6y ago
This kind of matches the performance I recorded when numpy is linked to vecLib for large matrix matrix multiplication in float32:
https://gist.github.com/ogrisel/87dcf2c3ab8a304ededf75934b116b61#gistcomment-3614885 https://gist.github.com/ogrisel/87dcf2c3ab8a304ededf75934b11...
Note however there is currently no way to build and link numpy and scipy against vecLib to get correct results when calling LAPACK routines (to get Singular Value Decomposition for instance). It might be related to low level fortran ABI problems but I am not an expert so I don't know for sure.
It's possible to get a fully working numpy / scipy stack with OpenBLAS and gfortran by using the conda-forge distribution:
https://github.com/conda-forge/miniforge#download https://github.com/conda-forge/miniforge#download
The performance is not as good as with vecLib (see the linked benchmark) but it's already very good (e.g. compared to a similarly priced Intel or AMD laptop with OpenBLAS and maybe even MKL).