3 ms·
The PyT version above and the both the CPU/GPU versions on Neantherdal work completely in float32. I confirm that the performance of Neantherdal GPU is similar
by akssri 7y ago
The PyT version above and the both the CPU/GPU versions on Neantherdal work completely in float32. I confirm that the performance of Neantherdal GPU is similar to that in the above PyT version. Yes.
The part that is flawed is that you're comparing this to Numpy (CPU)/ Cupy (GPU) both of which coerce the input array to float64 (for precision reasons) before computing the covariance and correlation. You only need to check the output type of the result to verify this (if the pointer to the code is not sufficient).
- dragandj 7y agoWhat is flawed there? The point of the article is to show that CuPy very often does not accelerate NumPy, especially on consumer-grade hardware. This is something that most users of NumPy/CuPy do not know, and they are led by the docs to think it does. The reason for that is that CuPy is poorly implemented. And CuPy is poorly implemented because it is constrained by what NumPy does, which, in turn, does stuff that is OK on the CPU, and often translates poorly to the GPU.
- shezi 7y agoI didn't get that from the article at all. There was no "that's because..." or "you need to do this to make it fast..." in the article. Instead, it's "clojure is faster without additional work". While that's super neat and thank you for showing this, bashing python because the library doesn't automatically do that and then hiding behind this silly argument isn't that enlightening.
- dragandj 7y agoYou can check in the article that I made many checks to make sure that NumPy/CuPy get the data in float32. What do you suggest to do to instruct NumPy/CuPy to "automatically do that". Is there a way to say to python "I want you to use float32 precision for this computation" other than, well, providing everything as float32? And even if your argument stands, that does not change the fact that CuPy does not accelerate NumPy (in this particular case, but I'd say often).
- didibus 7y agoI didn't read it as bashing Python, only as bashing NumPy/CuPy. If you look at the article, it is not even using Python anywhere, it uses NumPy/CuPy directly from Clojure. So it really is comparing Neanderthal vs NumPy/CuPy, and uses Clojure in both cases.
- akssri 7y ago> What is /flawed/ there? You're comparing float32 vs float64 computation. I don't need to tell you how much slower DGEMM is vs SGEMM esp. on the GPU (you mention this in the post yourself!). Numpy does this for precision reasons, and CuPy simply follows its behavior. This is precisely why I noted that the float32 version runs 3x faster on the CPU. > The reason for that is that CuPy is poorly implemented. It's a cheap shot to call something 'poorly implemented' when you don't understand what you're benchmarking.
- dragandj 7y agoCan you please provide a benchmark for corrcoef where CuPy is noticeably faster than NumPy on your (and mine) GPU, Nvidia GTX 1080Ti?
- didibus 7y agoI still think I'd have to disagree with you. What was benchmarked was NumPy/CuPy, and the numbers in the article are not flawed. It isn't that they are using NumPy/CuPy wrongly, that's what you'd do, and even if you try really hard to specify everything as float32 it still will have the same performance timing as in the article. It would be interesting to compare it against a float64 version in Neanderthal as well I agree with that. That said, a flawed benchmark would mean to me that it isn't indicative of the performance one can expect when actually using the library on real world use case, but for now this benchmark for NumPy/CuPy does seem to be indicative of what you'd expect. Now, the next question is, for model accuracy vs scale, is going with float64 coercion always the ideal trade off? What if you still needed to squeeze more performance? Is it really a bad idea to do so by going down to float32? Especially considering how much faster GPU can accelerate that?
- bearzoo 7y agoThe flaw, in my opinion, is not necessarily in the benchmark but in the claims and observations around the benchmark. For instance the blog flat out says: > Nope, we work with float32. It seems that is not true. The blog should make clear that it isn't straightforward or possible to get cupy to do this.
- 7y ago
- rrss 7y ago> I would love to improve any part of this article, if possible! You should state that the reason the neanderthal version runs faster is that it is doing the computation in lower precision than numpy / cupy. The article walks through an investigation of why cupy's result is underwhelming (is the input data accidentally fp64? is the input data on the cpu? is the computation happening on the cpu?), so you should finish it by explaining that numpy and cupy do the computation in fp64.