2 ms·
> You can't just blindly compare flops numbers like this. Hitting 13tflops on a GPU implies data has long been loaded into GPU global memory and tons of computa
by leeter 8y ago
> You can't just blindly compare flops numbers like this. Hitting 13tflops on a GPU implies data has long been loaded into GPU global memory and tons of computations happen on it.
True, saturating a GPU is actually really really hard.
> Not to mention you're limited to 16GB at a time on a Radeon VII and you'd need to either use OpenCL or RoCM which is time consuming.
On a consumer card yes... on professional cards 32GB and up is common. That said most workloads aren't THAT big, and even if they are you can usually stage it in such that the GPU isn't ever really idle.
Vulkan and DX11+ support compute shaders now. MSVC has a custom extension too. There are plenty of compute libraries that exist that will prebuild and do all of this for you. It took me half an hour max to get set up to do this, there are reasonably good tutorials out there.
> Fact is, GPUs are great for a minority of workloads (definitely not most) that are very computationally intense but not memory heavy. As soon as you have lots of reductions, etc. where threads will wait a CPU will also be a much better choice.
Going to disagree here, memory bandwidth on a GPU is insane. HBM2 absolutely blows GDDR5 out of the water on total bandwidth and that blows regular old DDR4 out of the water as well. This is because GPUs would be dead on arrival if they can't actually process data that fast, in fact it's arguable that GPUs are more memory bandwidth constrained than they are anything else. This explains why AMD spent so much of the budget on the Radeon VII on HBM2.
As for reductions, what do you think Anti-aliasing is? There are very well known tile based approaches for this that again, blow a CPU out of the water.
For a GPU the biggest holdback is actually dataset size. Too small and the cost to go over the PCI-E bus in terms of latency will override the savings. If you can do the entire calculation as a series of shaders then write the result back to main memory you can easily out perform any CPU on the market (there are a lot of caveats here however in regards to datatypes, but let it suffice that a professional level card will).
- dragontamer 8y ago> True, saturating a GPU is actually really really hard. That's true, but GPUs have over 100x the FLOPs of a CPU. This means that if you write a GPU-program with only 5% utilization (that's 95% idle), you still are going to be 5x faster than a perfectly utilized CPU. When you have 100x the raw FLOPs and 10x the memory bandwidth... you can write a program that only utilizes 15% of a GPU and still end up with results way faster than any CPU would give you. -------- CPUs are way easier to program though. CPUs actively look for work to do on your behalf: its like there's an optimization engine built into the chip itself (out-of-order scheduler, branch prediction, etc. etc.). So achieving high-utilization on CPUs is an easier job (but still hard).