4 ms·
This is pretty much the analogue of gather vs scatter in GPGPU programming. It's a well known fact that in GPU programming, in almost any case a gather approach
by w0utert 6y ago
This is pretty much the analogue of gather vs scatter in GPGPU programming. It's a well known fact that in GPU programming, in almost any case a gather approach (threads map to outputs) works better than a scatter approach (threads map to inputs). The challenge is to transform the algorithm to still have some data locality for the reads to allow for caching, coalesced reads into local memory etc, which can be very hard or infeasible.