5 ms·
This is a great write-up. I think Rust could be a very nice language for Shader-like applications which are already very functional, and don't involve alot of s
by AndrewGaspar 7y ago
This is a great write-up. I think Rust could be a very nice language for Shader-like applications which are already very functional, and don't involve alot of shared, mutable state across threads.
In HPC, we're very much interested in GPU compute programming, rather than shader programming. In CUDA codes, you're typically doing transformation from input buffers directly into output buffers from your CUDA kernels. This should immediately raise red flags for a Rust developer - you've got shared, mutable state across threads!
Consider this simple CUDA-ish Rust code with threads independently executing over 0..cuda.len() (ignore the bounds bugs at i = 0 and i = in.len()):
fn stencil(i: usize, in: &[f32], out: &mut [f32]) {
out[i] = (in[i - 1] + in[i] + in[i + 1]) / 3.0;
}
(The `i` is a conceit around computing indexing from thread/block IDs, but the input and output arrays are pretty similar to the style CUDA promotes).
It's obvious to me, the programmer, that I don't have any aliasing issue - each thread is only mutating at a single index in the output array. However, Rust is not smart enough to see this. If they allowed the definition of the kernel as is, you could easily write multi-threaded code that has shared mutable access to individual memory locations, violating Rust's memory model. OK, you force the kernel to look more like this, then:
// `in` is the slice of [i-1,i+1]
fn stencil(in: &[f32], out: &mut f32) {
*out = (in[0] + in[1] + in[2]) / 3.0;
}
And you'd enforce Rust's invariants at the kernel launch site, computing the valid slices at some higher level in the library in some "unsafe" code. But this only solves the simple case where you have some array mapping to another array where the index relationship is obvious, and it's easily provable that there are no aliasing issues. Start layering in things like unique indirect indexing, or perhaps non-unique indexing but with atomic reductions, and it becomes difficult to phrase your correct program in a way to safe(!) Rust that is compatible with the borrow checker, at least without having to build a bunch of abstractions to express each of your parallel patterns. Having to build a bunch of bespoke abstractions may not be scalable to the types of developers building big scientific codes.
Anyway, I'm curious if the folks at "Embark" have spent any time thinking about the issue of shared, mutable state in GPU programming with Rust. It seems like a deal breaker from where I stand.
- gmueckl 7y agoYour point is especially important as most of the GPGPU performance gains come from really clever use of shared memory. That is the very model of shared state with concurrent acces. Usually, your shader or kernel is required to synchronize threads that need to see each other's changes. I'm sure that there are at least some kernels out there that skip synchronization and still work because there are enough other instructions between write and read accesses that cover the issue up.
- dragontamer 7y ago> shared memory. I've done enough Hacker News discussions to note that the lay-reader won't understand what this means, and this probably needs more elaboration. "Shared Memory" is a special memory area inside of GPUs where grids (NVidia) or workgroups (AMD)... a group of 32x to 1024x SIMD threads... can perform inter-thread communications in an outrageously fast way. Shared memory is extremely small: roughly 64kB in size. Optimizing shared memory access involves resolving bank-conflicts, and lots of very-low level thought. At a minimum, you need to consider how you fill shared memory (the memcpy in) as well as how to get the final data out (the memcpy out). --------- In many cases, synchronization to-and-from shared memory only requires a __threadfence() instruction. Maybe only a __syncwarp() instruction in some cases. A lot of thought goes into the "ordering" of memory accesses, to make shared-memory as fast as possible. See here for further details: https://developer.nvidia.com/gpugems/GPUGems3/gpugems3_ch39.html https://developer.nvidia.com/gpugems/GPUGems3/gpugems3_ch39.... See "39.2.3 Avoiding Bank Conflicts" for how shared memory is optimized in practice. You're not only thinking about the shared-state, but also the average number of other threads hitting any particular bank. On a say... 32-bank system, you'll want all 32-banks to be utilized as much as possible. If all 1024 threads are accessing bank#0, you'll be 1/32th the speed (and banks #1 through #31 are all wasted). A "bank" is basically the implicitly "RAID0" arrangement of shared-memory. The precise number of banks is somewhere between 16x banks or 32x banks, depending on architecture. But regardless, if all GPU-threads access memory location #0, that will all hit bank#0, which is slow. Instead, you want GPU-threads to "spread out" over the banks (maybe GPU-thread#0 should access Bank0 / Memory location #0. GPU-thread#1 should access bank1 / memory-location #16. GPU thread #2 should access bank2 / memory-location #32. Etc. etc.)
- pcwalton 7y agoThere's no memory safety problem if the data you're racing on is pointer-free. Rust allows racing on memory with Relaxed atomics [1]. (Yes, I know about potential UB with floating point and whatnot; this is solvable.) Happily, GPU programming tends to use a lot of indices as opposed to direct pointers, because it makes CPU/GPU memory management easier--indices are valid no matter where in the address space the data in question is mapped. Of course, you might then ask "what's the point of using Rust at all?" The answer is that a lot of code can be expressed in regular patterns idiomatic to the Rust type system and borrow checker. This isn't particularly different from any other type of programming. You sometimes see people say "Rust is useless because you can't safely write an arbitrary graph with direct pointers". The obvious fallacy with this argument is that most code doesn't need an arbitrary graph with direct pointers; for the small amount of code that does, you can drop into unsafe in a targeted way and still have a much safer program than one written in e.g. C. GPU programming is no different. [1] https://www.cl.cam.ac.uk/~pes20/cpp/cpp0xmappings.html https://www.cl.cam.ac.uk/~pes20/cpp/cpp0xmappings.html (Note than on x86 Relaxed atomics turn into plain mov instructions, and on ARM they turn into plain ldr and str.)
- AndrewGaspar 7y agoHuh - to make sure I understand correctly, is it your view (ignoring portability concerns) that it should be legal to store and load directly from a non-mut &[f32] when the platform has relaxed-atomic memory consistency by default? Or just that you should use a hypothetical "AtomicF32" with relaxed loads and stores and not worry about it from a performance perspective?
- pcwalton 7y agoThere should be an AtomicF32 type that supports unsynchronized reads/writes on GPU.
- AndrewGaspar 7y agoThanks, that makes a lot of sense!