4 ms·
GPUs can hide memory latency very well, because they are basically SMT on steroids. (Imagine instead of 2 threads per "core" you have 32-64 threads) But they a
by PixelOfDeath 6y ago
GPUs can hide memory latency very well, because they are basically SMT on steroids. (Imagine instead of 2 threads per "core" you have 32-64 threads)
But they are starved of memory bandwidth! And the lower latency memory CPUs prefer is not the same as the high bandwidth memory that GPUs like.
Also there is different kind of caching. Modern APUs often have two ways to access memory. Over there own cache, or over the CPUs cache. So shared memory for CPU<->GPU gets the full advantage of a cache, but still it is a trade off.
If you want some work done by the GPU part of an APU, by sharing a pointer, you can do that today. But from the point of view of the CPU there is no prediction beyond the "GPU do X" commands. And a very high latency until the job is done. So you need a minimum GPU job size for it to make sense.