3 ms·
I like how the broadcasting comes out in the end in Julia. Besides that Julia has quite many niceties that I could only wish Spiral had and I do think that they
by abstractcontrol 8y ago
I like how the broadcasting comes out in the end in Julia. Besides that Julia has quite many niceties that I could only wish Spiral had and I do think that they matter significantly even if they are not vital. This is definitely praiseworthy work.
Since I haven't done so and I really should have, if you or anyone else still looking at this thread are curious how GPU programming without macros looks like then take a peek at this:
https://github.com/mrakgr/The-Spiral-Language/blob/3a94ca644160310cb494c92ebf8c6a5ed9836f5f/Learning/CudaModule.fs#L1289 https://github.com/mrakgr/The-Spiral-Language/blob/3a94ca644...
Comparing Spiral and Julia code that has been linked so far, I feel that no one would ever guess without knowing ahead of time that Spiral is a static language and Julia dynamic based on these examples.
As an example of this in action, here is how layer normalization is implemented using the 'seq_broadcast' kernel which does a sequence of reductions and broadcasts in registers. It is also used to implement softmax for example.
https://github.com/mrakgr/The-Spiral-Language/blob/3a94ca644160310cb494c92ebf8c6a5ed9836f5f/Learning/LearningModule.fs#L171 https://github.com/mrakgr/The-Spiral-Language/blob/3a94ca644...
- simondanisch 8y ago>GPU programming without macros looks like then take a peek at this: I'm not sure if you missed my examples, or if there is some other missunderstanding going on ;) GPU code in Julia works 100% without macros, but you are free to use them, so people do ;) Also, you can pretty much write the gpu kernels like pseydo code, not sure how much simpler it can get. With your linked spiral code, I can't even really tell where the algorithm starts and where the setup code ends - which of course might easily be attributed to my unfamiliarity with Spiral ;) I wrote an article some time ago about generic programming with Julia's gpu infrastructure: https://techburst.io/writing-extendable-and-hardware-agnostic-gpu-libraries-b21c145a8dad https://techburst.io/writing-extendable-and-hardware-agnosti... If you can make concrete examples of how things can be simpler than this, I'd be delighted to hear them :) Cuda only (for no reason actually) and with lots of macros, here is another article about gpu programing in Julia: https://mikeinnes.github.io/2017/08/24/cudanative.html https://mikeinnes.github.io/2017/08/24/cudanative.html Do you have any benchmarks for the softmax kernel? If that kernel has optimal performance, it would be quite interesting. If it's sup par, it looks much longer than a simple version.
- abstractcontrol 8y ago> With your linked spiral code, I can't even really tell where the algorithm starts and where the setup code ends - which of course might easily be attributed to my unfamiliarity with Spiral ;) Heh, I've noticed that the setup code tends to come out longer than the actual kernels. It is inside `kernel = cuda`. Spiral is indentation sensitive like Python and F#. The stuff inside {} is just module creation, think of it like tuples with named fields. > Do you have any benchmarks for the softmax kernel? If that kernel has optimal performance, it would be quite interesting. If it's sup par, it looks much longer than a simple version. No, I've yet to actually benchmark it. It really depends on how good of a job NVCC does with the generic sequential reduce kernel. I'll do an in depth analysis when I am done with all the neural network work that I am doing currently which might take a while. https://github.com/mrakgr/The-Spiral-Language/blob/7ecd30bdf36e8f59dd89917c64a05794e2b5def1/Learning/LearningTests.fs#L431 https://github.com/mrakgr/The-Spiral-Language/blob/7ecd30bdf... You can see how it is implement here for the forward and backward parts. https://github.com/mrakgr/The-Spiral-Language/blob/7ecd30bdf36e8f59dd89917c64a05794e2b5def1/Learning/LearningModule.fs#L441 https://github.com/mrakgr/The-Spiral-Language/blob/7ecd30bdf... In the actual cost function, it is a bit different since I fuse the forward and the backward parts. It is a bit ugly because of all the type checking for whether the argument is a dual, but that is all done at compile time. > https://mikeinnes.github.io/2017/08/24/cudanative.html https://mikeinnes.github.io/2017/08/24/cudanative.html This is actually loop unrolling example that I had in mind when I said that Julia needs macros for. https://github.com/mrakgr/The-Spiral-Language#3-loops-and-arrays https://github.com/mrakgr/The-Spiral-Language#3-loops-and-ar... You can in fact get loop unrolling with just standard functions in Spiral. I go into it in the context of this chapter. It is achieved by pattern matching over tuples and recursion. > If you can make concrete examples of how things can be simpler than this, I'd be delighted to hear them :) Yes, by replacing meta programming with intensional polymorphism and inlinining guarantees. Also reflection should be done using pattern matching. I think this last one could be taken entirely seriously as it would not involve a replacement of the entire type system. All the claims in the article about inlining and specialization that make it sound like magic is what in general makes me dubious about languages pretending to be speed kings. Yes, I am aware that GPU kernels do not require optimizers capable of having monads for breakfast and that in the context of GPU programming where they were made they are probably true, but inlining is the sort of thing that matters more the more high level a language is. For very high level languages that desire speed, there isn't much choice but to make them a part of language semantics despite the added burden it puts on the user. Since we are still at it, I have a question I need to ask. Recently I've been informed that Julia is capable of GCing GPU memory. If this is fully integrated that would be a major feature which is not possible in say .NET or Racket. I really wanted this in Spiral and could not get it in .NET. By fully integrated, I mean much like for regular memory for which the GC takes note of the state of the system for when to do collection and defragmentation. If it can only make a thin wrapper with a finalizer (much like in .NET) around an unmanaged resource then it is not a big deal. Is it fully integrated or is it a wrapper style memory management? If it is the later, then that is too bad, but I'd suggest to Julia devs that they work on making it fully integrated as it would be a really good feature. Obviously, I can't do it in Spiral as I would need to write my own VM and I have only so many years in my life.