4 ms·
Outdated. Learn VULKAN
by keepquestioning 4y ago
Outdated. Learn VULKAN
- synergy20 4y agoare all big players supporting it? is there a good document. while this is all about graphics, I wonder how GPGPU will end up in the future. OpenCL did not get much love, CUDA still rules the market(leaves Intel's OneAPI and AMD's ROCm and the legendary OpenMP etc in the dust). Apple's Metal does both graphics and GPGPU, Vulkan is also said to do both graphics and GPGPU, not sure about Microsoft's, does it even have a GPGPU API considering it is the only player without its own hardware.
- psychphysic 4y agoWhat really killed them (as in not CUDA) is no easy way to use them. OpenCl pretty much broken on mesa so really amd is no better than Nvidia on Linux. And Nvidia killing it with CUDA where it works.
- raphlinus 4y agoDirectX12 has ok compute capabilities, but is increasingly being left in the dust by Vulkan. It doesn't have pointers or a memory model, but it does have device scoped barriers (which Metal lacks), so you can do things like single-pass prefix sum (monoid scan) operations, which Metal and WebGPU can't. Also, the Nvidia "cooperative matrix" operations (which access the "tensor cores") exist in CUDA and Vulkan but not DirectX12. Programming GPU compute outside CUDA is still quite painful; the developer experience of tools is terrible and the ecosystem has a lot of catching up to do. I'm most bullish on WebGPU going forward, but there are definite growing pains.
- dagmx 4y agoCould you elaborate on how device scoped barriers are benefficial? I’m not sure I follow where that would help for your example, but I’m definitely curious (and Google searches have become increasingly useless of late)
- raphlinus 4y agoSure! I have blogged about this[1], but I'll summarize here. Basically a barrier helps you do a "message passing" pattern, where one workgroup prepares some data then sets a flag, and another workgroup can read the flag and then read the data from the first workgroup. However, in Metal (and hence WebGPU) the barriers aren't powerful enough to guarantee you won't see stale data. Thus, to do prefix sum on Metal, you need to at least two dispatches, one to aggregate reductions over partitions, then another to do the sum within a partition. That's less efficient, both because of the dispatch overhead, and also because you need to read the data at least twice. You also need a bunch of different versions of your shaders to handle different problem sizes, which is really annoying (in fact, Vello can't handle more than 64k path segments until I write the larger version). Vulkan, CUDA, and DirectX12 (even DirectX11) can all do this; it's one of the ways in which Metal is an inferior basis for doing GPU compute. Btw, this is one of the reasons I feel a bit burned by MoltenVK, as it happily and silently translates correct SPIR-V into MSL that's lacking the correct barriers. In my experience, GPU translation layers are some of the leakiest abstractions around. [1]: https://raphlinus.github.io/gpu/2021/11/17/prefix-sum-portable.html https://raphlinus.github.io/gpu/2021/11/17/prefix-sum-portab...
- dagmx 4y agoThanks for the great write up!
- pjmlp 4y agoVulkan Compute hardly matters and DirectX future is on mesh shaders. Also all major graphics vendors tend to create their hardware designs in collaboration with Microsoft as part of DirectX, and eventually add them as extensions to Khronos APIs. So it hardly matters what Khronos is doing with Vulkan.
- tmtvl 4y agoall major graphics vendors tend to create their hardware designs in collaboration with Microsoft as part of DirectX That's pretty sickening, I hope the EU will give them a rap across the knuckles for that. Very anti-competitive.
- pjmlp 4y agoIt is up to Khronos to make it more interesting to do otherwise. Hardware Transformation and Lighting, CG/HLSL, RayTracing, MeshShaders and more recently DirectStorage, are all examples of features that appeared first in DirectX, with AMD or NVidia hardware implementation, before showing up as OpenGL or Vulkan extension.
- midnightclubbed 4y agoVulcan is on windows, Linux, Android, and OSX via moltenvk. Compute (GPGPU) is a requirement for all Vulcan drivers, technically drivers do not have to support rasterization but I don’t know of any hardware that doesn’t. Outside of Apple every other current maker of GPU hardware have native Vulcan drivers for at least some of their target products.
- pjmlp 4y agoMoltenVK acts as middleware and cannot support everything that isn't available on Metal. Outside of Apple there is zero support on XBox and PlayStation.
- cshenton 4y agoI personally find the experience of writing GPU compute code pretty nice on graphics APIs. The interface is pretty much the same “dispatch a 1-3D set of 1-3D work group indices”. The main pain points vs dedicated compute stuff like cuda is libraries and boilerplate to manage memory and launch kernels.
- synergy20 4y agoand, how to make the kernel and memory-allocation code working with tensorflow/pytorch, GPGPU is really now just a few libraries made for Tensorflow and Pytorch to invoke, same as CUDA, as far as ML is concerned.
- dahart 4y ago> I wonder how GPGPU will end up in the future. Didn’t the acronym GPGPU refer to when GPUs and APIs didn’t provide general compute capability? I thought this abbreviation was mostly out of use already, since we have CUDA and compute shaders now. It seems Metal and Vulkan and DirectX and OpenGL all use the term “compute” (Microsoft does have DirectCompute and compute shaders, they’ve been around for at least a decade). Not sure if that answers the question - all the graphics APIs provide a mechanism for compute shaders - or if you are asking more about a direct competitor for CUDA or something else?