5 ms·
One of the things I never understood about the hype around these Clojure GPU libraries is that there is a lot of marketing jumbo about how high-level it all is.
by hellofunk 6y ago
One of the things I never understood about the hype around these Clojure GPU libraries is that there is a lot of marketing jumbo about how high-level it all is. Even the book for these libraries has the words "no C++!" half a dozen times on its ordering page [0].
However, these are just high-level libraries and you cannot write the GPU shaders in Clojure. You still must use C or C++ for that (depending on if you are going the CL or Cuda route), and your C/C++ code is embedded or called from the Clojure side. This is no different than using another high-level language like Python -- I mean it's focusing on a rather misleading and irrelevant detail. You can use Clojure or Python (or many other languages) to call OpenCL or Cuda shaders, or you can just use the libraries that call them for you. Clojure is not special in this regard.
[0] https://aiprobook.com/deep-learning-for-programmers/ https://aiprobook.com/deep-learning-for-programmers/
- dragandj 6y agoThe book you've linked contains exactly 0 lines of C or C++ code. 100% of the code in the book is Clojure.
- didibus 6y agoHum.. I'm not sure there's that much hype. It's just saying that Clojure can also be used as a high level layer for fast linear algebra. I do see some claims that as a high level layer, Clojure works even better, it gives you a full REPL, and the power of macros and all that. It also goes to show you that there's barely any overhead added by the Clojure layer, which is something you want from a higher level layer. Maybe you need some additional context, but in the Clojure world, there are many offerings for that layering, core.matrix, neanderthal, NumPy, CuPy and others can all be used. So I guess a lot of what draganj is trying to show is he believes neanderthal is the layer with the least overhead, which everyone can agree to disagree on. You're never going to beat an assembly level implementation, or 50 year's worth of optimization done to a Fortran implementation, or straight openCL/cuda implementation in a higher level language, especially not in Clojure. So in the market of high level languages like Julia, Clojure, Python, R, Swift, etc. and their corresponding libraries, they all are just competing on being more user friendly, having better more useful/flexible abstractions, a better development flow, and the least amount of overhead. So that's what you'll be hearing them all fight over.
- akssri 6y agoCommon Lispers have had the ability to write CUDA kernels (with shaders) on-the-fly with a DSL in cl-cuda for many years now. See, https://github.com/takagi/cl-cuda https://github.com/takagi/cl-cuda Cupy (a project in which takagi is/was actively involved) also has the ability to compile kernels in the Python REPL, but the kernel code needs to be written in C and passed in as a string (pyOpencl can do the same for OpenCL). In fact, not too many years ago, the kernels for gradient descent, max pooling etc. were all entirely in Python in Chainer. Projects like JAX take this to another level by having the ability to transform the bytecode of a restricted class of Python functions straight into CUDA kernels. I can't speak much about Clojure/Neanderthal, but I'd advise the people to stop dissing projects they are not familiar with, least of all using silly benchmarks like the above.
- dragandj 6y agoGood advice. And, yet, you've taken it pretty seriously to diss Clojure/Neanderthal and my blog post, mostly by talking about unrelated stuff and projects that the post didn't even mention. And while introducing these themes left and right you didn't even bother to show some code related to this topic, just a suggestion of great projects by cool people. Yes, you showed the PyTorch code related to the blog post that confirms what the blog post says, but when I pointed out that the code has incorrect functionality (by missing some calculations) you didn't even bother to correct it, or to confirm that the code is good and that I'm wrong. So, it seems that your standard is that it is enough for one side to throw bits and pieces around and call it a day, and for the other to run around and prove that their stuff is better than everything that could possibly be done in every technology. I choose to stick to the theme. The theme is CuPy, NumPy, Clojure & Neanderthal. The related theme could be code in another technology. Great - write about it. But, even if every other technology were a million times better than what I describe in the article, it does not change the fact that CuPy and NumPy have the issue I've described.
- akssri 6y ago> And, yet, you've taken it pretty seriously to diss Clojure/Neanderthal and my blog post I have not - all I've said so far is that your benchmark is flawed. The fact that the code fragment above assumes zero mean data (thus using 2 fewer L1 ops) doesn't change a single thing in anything that has been written; to wit, the timings change to 28.6ms (GPU) and 333 ms (CPU). Pedantry is not an argument.