3 ms·
I've been playing around with some low-level GPU stuff lately and trying to port some of my CPU-bound matrix multiplication. The challenge has been that in the
by billti 2y ago
I've been playing around with some low-level GPU stuff lately and trying to port some of my CPU-bound matrix multiplication. The challenge has been that in the space I work in (quantum) the matrices are of complex numbers, and for a certain size of problem floats don't cut it and you need doubles. Nearly everything I find for GPUs is targeted at matrices of real floats.
Any pointers to examples? I'd be fine sticking with floats as a first step, but would love to see some (reasonably optimized) low-level GPU code for working with matrices of complex numbers (preferably for Metal or WGPU, which is what I'm using).
- LegNeato 2y agoCUDA has support for complex numbers and some NVIDIA frameworks use and expose it (https://nvidia.github.io/cccl/thrust/api/structthrust_1_1complex.html#_CPPv4I0EN6thrust7complexE https://nvidia.github.io/cccl/thrust/api/structthrust_1_1com...). For Rust GPU, nothing built in but there are libraries like https://github.com/rust-num/num-complex https://github.com/rust-num/num-complex that support `no_std` and should work on the GPU. I've never used them so I don't know what (if any) the perf hit would be.
- pythomancer 2y agoIn BLAS terminology this is usually called CGEMM (for single precision) or ZGEMM (for double precision). Both cuBLAS and rocBLAS support ZGEMM, and the latter is open source. rocBLAS is also pretty complicated though, and perhaps not such a good learning resource. This is a more readable library which implements CGEMM, or at least a similar operation: https://git.astron.nl/RD/recruit/ccglib https://git.astron.nl/RD/recruit/ccglib. The main issue is that double precision is not so interesting for AI and graphics, and so silicon is rather spent on more of these features and less double precision. Not so for HPC, though, and GPUs specialized for this usually have better throughput. For example, the AMD MI210 has the same performance in single and double precision (matrix) operations, while graphics GPUs either have something like 1/2, 1/4, 1/16 etc rate of fp64:fp16, or have no support at all.