3 ms·
It doesn't seem to me like you've described concepts that are all that different, loading from global memory is slow on the CPU too because often it requires so
by MintPaw 19d ago
It doesn't seem to me like you've described concepts that are all that different, loading from global memory is slow on the CPU too because often it requires some kind of barrier or lock. All techniques require you to pre-arrange memory, just in different sized blocks. And sharing data between "lanes" is always to be minimized. Different amount on different platforms ofc, but the ideas seem similar.