3 ms·
I benchmarked some of my large transformer networks the last few days and MKL is still 50% faster than OpenBLAS. What's even worse in real-world applications i
by danieldk 6y ago
I benchmarked some of my large transformer networks the last few days and MKL is still 50% faster than OpenBLAS.
What's even worse in real-world applications is that OpenBLAS misbehaves when an application uses threads. This is also described in the OpenBLAS FAQ:
If your application is already multi-threaded, it will conflict with OpenBLAS multi-threading. Thus, you must set OpenBLAS to use single thread as following.
https://github.com/xianyi/OpenBLAS/wiki/faq#multi-threaded https://github.com/xianyi/OpenBLAS/wiki/faq#multi-threaded
- gnufx 6y agoSo what are the results with libxsmm and current AMD BLAS, as that must be for small dimensions? The reason it's serial BLAS that mainly matters is that HPC codes are usually parallelized above the BLAS; why do you want the nesting? Swapping in threaded OpenBLAS or BLIS is something you might do with basically serial stuff like vanilla R, e.g. https://loveshack.fedorapeople.org/blas-subversion.html#_addendum_example https://loveshack.fedorapeople.org/blas-subversion.html#_add... OpenBLAS threading has been somewhat buggy, but the main problem with its OpenMP support currently seems to be that using OMP_PLACES kills it.