4 ms·
Dense matrix algebra has the benefit that besides the matrix dimensions, the data itself has no impact on performance I was very surprised when I first learned
by imurray 8y ago
Dense matrix algebra has the benefit that besides the matrix dimensions, the data itself has no impact on performance
I was very surprised when I first learned that's not strictly true, because of denormal numbers[1]. Here's a session in ipython --pylab (with MKL), demonstrating a 200x slow-down for matrix multiplication with tiny numbers. Crazy!
In [1]: A = randn(1000, 1000)
In [2]: %timeit A @ A
25.9 ms ± 199 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)
In [3]: A *= 1e-160
In [4]: %timeit A @ A
5.52 s ± 21.2 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
You hit denormal numbers more quickly with single-precision floats. I have been bitten by this issue in the wild a couple of times now, and seen a couple of other people with it too. Sometimes denormals are created internally in algorithms, when you didn't think your input matrices were that small.
[1] https://en.wikipedia.org/wiki/Denormal_number#Performance_issues https://en.wikipedia.org/wiki/Denormal_number#Performance_is...
- CamperBob2 8y agoIt's just incredible that there's still no way to force denormals to be rounded silently to zero on x86. You have to go out of your way to check for them, or face horrific performance penalties for a feature you didn't even want or need.
- imurray 8y agoThis code: #define CSR_FLUSH_TO_ZERO (1 << 15) unsigned csr = __builtin_ia32_stmxcsr(); csr |= CSR_FLUSH_TO_ZERO; __builtin_ia32_ldmxcsr(csr); from https://stackoverflow.com/a/8217313/ https://stackoverflow.com/a/8217313/ seems to work for me with gcc. I don't know how widely supported it would be.
- stephencanon 8y agoHuh? Set bits 6 (DAZ) and 15 (FZ) in MXCSR. Done. The only instructions you can’t set to flush are the legacy x87 opcodes, which you shouldn’t be using in a performance-sensitive context anyway.
- CamperBob2 8y agoSadly, some of us are stuck maintaining code that still needs to run on pre-SSE2 hardware.
- modeless 8y agoYes! For some applications including machine learning you can set the option to flush denormals to zero so you don't hit this case. Infinity and NaN can also be a performance problem but it depends on your processor. Though of course if you get those during matrix multiplication something has probably gone wrong.
- adrianratnapala 8y agoDepends on what you mean by wrong. The point of NaNs and infinities is that sometimes it is easier/faster/cleaner to just let those things propagate through the calculation than to do conditional logic that in the end just tends to just simulate the same effect. True, that consideration is more valid for element-by-element operations rather than things like matrix multiplications. But depending on the operation it might make sense to rely on NaN propagation rather than trying to prescan you matrix for NaNs.
- dekhn 8y agoThe quote was referring to the distribution of the data in the matrix, not the specific details of data in the matrix (such as denormals).