4 ms·
A few thoughts: * This might be a good candidate for a compute shader, the swizzle abilities of shaders make quite a few of the things they are doing much easi
by leeter 8y ago
A few thoughts:
* This might be a good candidate for a compute shader, the swizzle abilities of shaders make quite a few of the things they are doing much easier as does the faster memory connections GPUs (usually) have.
* The committee is actually looking at potential ways to handle this http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p1101r0.html http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p110... has been proposed. There is also a lot of discussion about support for GPUs in a standard way, but I haven't seen any proposals myself.
- gameswithgo 8y agoiirc the domain this guy is working in, is one where the GPU is going to be busy with rendering, so he may prefer to keep this on the cpu. But you are right that any time you turn to SIMD, you might be able to turn to the GPU and get a much biggger win.
- leeter 8y agoSo what I'm somewhat curious about is if this isn't an argument for heterogeneous GPU usage. Schedule this on an iGPU while the main GPU is busy. The main downside, as I alluded to above, would be memory bandwidth unless there is a good on die cache for the iGPU or it has dedicated memory. That said with a multi-gpu setup you might also have another card you could use. The other thought I have is that it's very difficult to keep the GPU 100% saturated in my experience. If the vertex count is high enough to make the PCI-e transition worth it (and it would seem to be), then you can potentially cycle steal a bit on the main GPU. This will potentially slow down the main render but speed up the overall. This would be particularly true if you can rebind the resource back for use in the main render pipe.
- gameswithgo 8y agoAll ideas that can work in certain situtions. As GPUs get more like CPUs and CPUs get more like GPUs things will just get fuzzier I imagine.
- dragontamer 8y agoA $850 CPU (Threadripper 2950x) will give you Hundreds of GFlops with ~100GB/s bandwidth to main memory. A $700 GPU (AMD Radeon VII) will give you 13.8 TFlops with 1TB/s bandwidth to video memory. If you have a workload that lends itself to GPU programming, then you should buy a 2nd GPU (or 3rd, or 4th GPU) and stick it into your system. Not all workloads work on a GPU, but a LOT of workloads just make more sense on a GPU. > But you are right that any time you turn to SIMD, you might be able to turn to the GPU and get a much biggger win. As long as the problem is big enough to mitigate the PCIe latency, kernel startup time, and PCIe bandwidth. Transferring data is relatively slow (PCIe x16 == 15GB/s). There are many algorithms which are faster to execute on the CPU, because executing out of L1 cache is faster than waiting for the PCIe bus. In these cases, you should prefer CPU-side SIMD, to keep the data "hot" in L1 cache.
- alfalfasprout 8y agoI bring this up everytime a comment like this comes up... You can't just blindly compare flops numbers like this. Hitting 13tflops on a GPU implies data has long been loaded into GPU global memory and tons of computations happen on it. Not to mention you're limited to 16GB at a time on a Radeon VII and you'd need to either use OpenCL or RoCM which is time consuming. Fact is, GPUs are great for a minority of workloads (definitely not most) that are very computationally intense but not memory heavy. As soon as you have lots of reductions, etc. where threads will wait a CPU will also be a much better choice.
- dragontamer 8y ago> As soon as you have lots of reductions, etc. where threads will wait a CPU will also be a much better choice. Two things: 1. Reductions have been implemented efficiently on GPUs for a while now. See this implementation of Parallel Prefix Scan https://developer.nvidia.com/gpugems/GPUGems3/gpugems3_ch39.html https://developer.nvidia.com/gpugems/GPUGems3/gpugems3_ch39.... That right there is your standard "reduce" algorithm on a GPU. Its work-efficient and utilization of the GPU remains very high. 2. GPUs have a ridiculously efficient "barrier" instruction. In fact, GPU "Barriers" are implemented at the assembly level on GPUs. Thread synchronization is actually very, very, very cheap on GPUs. In fact, if your workgroup matches the wavefront / warp of the underlying GPU, then a barrier is simply a NOP: as cheap as it gets. The main issue with GPUs is that they are a bit mysterious right now. Not a lot of people program them. But the more I work with them, the more I realize how flexible and fast their synchronization tools are.