4 ms·
Prefix sum on portable compute shaders
- dahart 5y ago> there is something slightly taboo about coordination between workgroups in the same dispatch. Many GPU experts I’ve talked with express skepticism that this can work at all. Even so, interest in this type of pattern is picking up, in part because of advanced rendering engines like Nanite, which also uses atomics to coordinate work between workgroups in a live-running dispatch. I live only in CUDA land, I’m a bit ignorant of the other platforms, so I’m not sure what granularity of workgroups is here. Generally speaking you can’t rely on individual threads executing in any particular order, no matter how many there are, so yes you can’t have one thread look at data from another thread without some kind of synchronization. It took me longer than it should have to really grok this and to believe it deeply, initially I kept wanting to think that thread number 235 million could read from thread 0 safely without synchronizing because surely enough time has passed. Nope, threads can and do come in any order, it’s never safe to assume another thread has finished. Using atomics to solve this is rarely a good idea, atomics will make things go slowly, and there is often a way to restructure the problem so that you can let threads read data from a previous dispatch, and break your pipeline into more dispatches if necessary. CUDA at least has a few other ways to share data during a single dispatch (or “launch”) besides atomics. Threads in warps can talk to each other, threads in blocks can share memory. But this all takes careful design and various other synchronization primitives. > That is sadly not the case for GPU compute code. The most common scenario is dependency on a large, vendor-dependent toolkit such as CUDA (that installer is a 2.4GB download). If you have the right hardware, and the runtime installed properly, then your code can run. While GPU coding is indeed more onerous than CPU programming in general, I feel like this wasn’t necessarily a fair point - this is the CUDA SDK download being compared to the CPU runtime. Installing the Rust compiler & cargo wasn’t mentioned as a downside, for example. CPU code also requires the right hardware and to have the runtime installed properly, it’s just something most people already have setup. Similarly, compiled CUDA code will run just fine without the SDK, and most people attempting to run a compiled CUDA program will have a compatible driver installed already. For the average compiled app, the run-time is no trickier than CPU, and the CPU has more or less the same kinds of requirements.
- CoolGuySteve 5y agoUnreal Engine's Nanite technology has the same problem. They want to walk a LoD/culling tree but have to resort to the CPU just to schedule new tasks on the GPU. I wouldn't be surprised if future GPUs eliminate this requirement as it will make all Nanite Unreal Engine games run a little faster. Kind of like how a couple years after Quake came out, every GPU got fast z-buffers. https://youtu.be/eviSykqSUUw?t=1616 https://youtu.be/eviSykqSUUw?t=1616
- raphlinus 5y ago> They want to walk a LoD/culling tree DAG, my friend, not tree. Seriously, I recommend people watch the talk, it's one of the more impressive demonstrations of how to use GPU compute power I've ever seen, and the results speak for themselves.
- raphlinus 5y agoWorkgroup in Vulkan/WebGPU lingo is equivalent to "thread block" in CUDA speak; see [1] for a decoder ring. > Using atomics to solve this is rarely a good idea, atomics will make things go slowly, and there is often a way to restructure the problem so that you can let threads read data from a previous dispatch, and break your pipeline into more dispatches if necessary. This depends on the exact workload, but I disagree. A multiple dispatch solution to prefix sum requires reading the input at least twice, while decoupled look-back is single pass. That's a 1.5x difference if you're memory saturated, which is a good assumption here. The Nanite talk (which I linked) showed a very similar result, for very similar reasons. They have a multi-dispatch approach to their adaptive LOD resolver, and it's about 25% slower than the one that uses atomics to manage the job queue. Thus, I think we can solidly conclud that atomics are an essential part of the toolkit for GPU compute. You do make an important distinction between runtime and development environment, and I should fix that, but there's still a point to be made. Most people doing machine learning work need a dev environment (or use Colab), even if they're theoretically just consuming GPU code that other people wrote. And if you do distribute a CUDA binary, it only runs on Nvidia. By contrast, my stuff is a 20-second "cargo build" and you can write your own GPU code with very minimal additional setup. [1]: https://github.com/googlefonts/compute-shader-101/blob/main/docs/glossary.md#workgroup-vulkan-webgpu-threadgroup-metal-thread-block-cuda https://github.com/googlefonts/compute-shader-101/blob/main/...
- 6d65 5y agoAnother great post in the series. I've been trying to bootstrap my own deep learning framework in Rust for a while now. I'm still stuck at implementing the stuff on the CPU. But I've always wanted something running on GPU as well. I've poked at OpenCL, and webgpu. But Piet-gpu(more probably piet-hal) seems to be the best starting point in the Rust land. Once I get to the GPU compute part, I'll start with Piet GPU. Not looking forward to debugging GPU compute kernels, but maybe it will be fun. PS. It's a bit disheartening to see such buggy Vulkan implementations, since one of the main Vulkan selling points was less bugs in the drivers. I'm not sure what I'll see on Linux. Also, I hoped that between Vulkan and Metal one could cover all the the major desktop OSes with GPU accelerated software. It's sad to see apple dropping the ball in this regard. There is also the concern of code reuse and code quality. Doing something complex in one pass would mean a huge kernel. It would be interesting to investigate if something like openai's Triton can be implemented in top of Piet-gpu(HAL). A dsl embedded in Rust that could do kernel fusion, and generate one shader when composed together. I'll poke around at this when I get to Piet-gpu.
- raphlinus 5y agoLet's talk if and when you want to build something. I make no promises that piet-gpu-hal is suitable for other workloads (right now I would say it barely meets the needs for piet-gpu), but on the other hand I am interested in the question of what it lacks and what it would take to run, eg, machine learning workloads on top of it.
- kvark 5y ago> Vulkan is defined in terms of the SPIR-V intermediate representation, so presumably can run anything that compiles to that “Presumably” doesn’t hold here, unfortunately. We are seeing issues in spirv processing by drivers if it doesn’t look like the output of glslc.
- masahi 5y agoTVM, similar to IREE, also has a good support for Vulkan. It compiles a tensor DSL written in Python into SPIR-V. AMD is using it in production for running deep learning models on their APU. https://github.com/apache/tvm https://github.com/apache/tvm