6 ms·
The benchmark is flawed. Cupy follows Numpy in doing a type-promotion of the data to float64 before computing the covariance. https://github.com/cupy/cupy/blob
by akssri 7y ago
The benchmark is flawed. Cupy follows Numpy in doing a type-promotion of the data to float64 before computing the covariance.
https://github.com/cupy/cupy/blob/b19577199d3f75c76e23c751941cb0976ce6ea01/cupy/statistics/correlation.py#L86 https://github.com/cupy/cupy/blob/b19577199d3f75c76e23c75194...
The following function in PyTorch completes the float32 computation in 27 ms (1080 Ti). I'm fairly sure that a similar function in Cupy would achieve very similar results.
def corr(dat):
cov = (dat @ dat.T).div_(dat.shape[-1]-1)
sigma = torch.diag(cov).sqrt_()
return cov.div_(sigma[None]).div_(sigma[None].T)
The heart of the computation SGEMM is in anycase provided by CUBLAS, and the only areas in which PyT and Cupy differ are in the tricks used to unroll the loop for the broadcasting & L1 op. The performance difference b/w PyT, Cupy and others should be negligible IMO.
Note:
np.corrcoef takes about 1.12 s and the above float32 version takes 325ms on the CPU.
- dragandj 7y agoSo, you first state that the benchmark is flawed, then you confirm the results of the benchmark, and repeat everything that the OP describes as the solution, just with a different tool (PyTorch instead of Clojure)? BTW did you check that your corr function gives the correct result? It looks to me that you forgot a few steps.
- akssri 7y agoThe PyT version above and the both the CPU/GPU versions on Neantherdal work completely in float32. I confirm that the performance of Neantherdal GPU is similar to that in the above PyT version. Yes. The part that is flawed is that you're comparing this to Numpy (CPU)/ Cupy (GPU) both of which coerce the input array to float64 (for precision reasons) before computing the covariance and correlation. You only need to check the output type of the result to verify this (if the pointer to the code is not sufficient).
- dragandj 7y agoWhat is flawed there? The point of the article is to show that CuPy very often does not accelerate NumPy, especially on consumer-grade hardware. This is something that most users of NumPy/CuPy do not know, and they are led by the docs to think it does. The reason for that is that CuPy is poorly implemented. And CuPy is poorly implemented because it is constrained by what NumPy does, which, in turn, does stuff that is OK on the CPU, and often translates poorly to the GPU.
- shezi 7y agoI didn't get that from the article at all. There was no "that's because..." or "you need to do this to make it fast..." in the article. Instead, it's "clojure is faster without additional work". While that's super neat and thank you for showing this, bashing python because the library doesn't automatically do that and then hiding behind this silly argument isn't that enlightening.
- dragandj 7y agoYou can check in the article that I made many checks to make sure that NumPy/CuPy get the data in float32. What do you suggest to do to instruct NumPy/CuPy to "automatically do that". Is there a way to say to python "I want you to use float32 precision for this computation" other than, well, providing everything as float32? And even if your argument stands, that does not change the fact that CuPy does not accelerate NumPy (in this particular case, but I'd say often).
- didibus 7y agoI didn't read it as bashing Python, only as bashing NumPy/CuPy. If you look at the article, it is not even using Python anywhere, it uses NumPy/CuPy directly from Clojure. So it really is comparing Neanderthal vs NumPy/CuPy, and uses Clojure in both cases.
- akssri 7y ago> What is /flawed/ there? You're comparing float32 vs float64 computation. I don't need to tell you how much slower DGEMM is vs SGEMM esp. on the GPU (you mention this in the post yourself!). Numpy does this for precision reasons, and CuPy simply follows its behavior. This is precisely why I noted that the float32 version runs 3x faster on the CPU. > The reason for that is that CuPy is poorly implemented. It's a cheap shot to call something 'poorly implemented' when you don't understand what you're benchmarking.
- jsdavo 7y agoThanks for sharing! Why does CuPy switch to float64 in this case? Is it a bug?
- akssri 7y agoNumpy/Cupy do explicit coercing to float64. There is no documentation for why this is done, but since GEMM is used to compute the covariates (summation over data points), it makes sense to increase the precision. You could get away using a double to only accumulate the sum, but it's a pain to write such mixed-precision (slow) function in C and then wrap it in Python, esp. for something like this. (See, https://github.com/numpy/numpy/blob/v1.17.0/numpy/lib/function_base.py#L2374 https://github.com/numpy/numpy/blob/v1.17.0/numpy/lib/functi...)
- fluffything 7y agoCuPy is a drop in replacement for NumPy. Ask Numpy why they do that.
- didibus 7y agoI'm not sure it is flawed. It just seems that you found the root cause of why NumPy/CuPy are slower, but it seems there is no way as a user of NumPy/CuPy to perform the float32 computation, even though you were explicit and said `:dtype "float32"`. All these libraries under the hood delegate to the same things for the actual computation, so the only advantage of one over the other would be this type of stuff. That NumPy/CuPy coerce to a float64 under the hood, and gives us no way to force it not to do so, is a disadvantage that they have in my opinion. In any case, like the article says, the performance of NumPy/CuPy is still good enough, and maybe NumPy thinks it knows better then it's users that they should be using higher precision always. Edit: I know I'll face the brigade of NumPy users. But I invite people to engage in the discussion unbiased. With a switch from CPU to GPU you'd expect more than 29% performance boost. If CuPy went with a float32 proc, it would give you a 96% performance increase! So ya, I think we can start to argue CuPy isn't doing the right thing here.
- akssri 7y ago> user of NumPy/CuPy to perform the float32 computation This is just getting tiring. Numpy and Cupy are perfectly capable of doing float32 computation - their only "fault" is that they coerce data to float64 in this one fairly unimportant function (which you can implement to your liking in 3-4 LOC). Hell, PFNs entire deep learning system, Chainer (which mind you, both predated and inspired PyTorch, and is still quite competitive with it), is built entirely on top of Cupy! Benchmarks make sense only when the outputs are the same - in this case, they certainly are not. It's the responsibility of the author to make sure that the outputs are the same (a real benchmark), or to argue that Numpy/Cupy are wasting time by using float64 in xp.cov (an issue of implementation). He does neither.
- didibus 7y ago> Numpy and Cupy are perfectly capable of doing float32 computation You're arguing that they would perfectly be capable of an implementation which would use float32, ya sure, if you change their source code you can have them be faster, that's an obvious to me? The question is, what kind of performance can you expect as a user using them as a library, and this benchmark is revealing of that in this case. What is wrong with it? It's like saying that you can change the implementation of Python and it will make Python code run faster, like sure you can, but you don't benchmark hypothetical future versions, the current version of CuPy uses float64 for this function, and that makes it that it barely runs faster then the CPU version. Point in case, I don't know what else you're trying to disagree about here. > their only fault is that they coerce data to float64 in this one fairly unimportant function I don't know that, all I know is when picking one function and benchmarking it, we found a flaw which result in performance gains from CPU to GPU to be extremely underwhelming. How many more if we started to benchmark the rest?