3 ms·
On the GPU, vectorization is done for you by the compiler, and the cores are designed to hide latency. These techniques will do little to improve that. OTOH, y
by superjan 3y ago
On the GPU, vectorization is done for you by the compiler, and the cores are designed to hide latency. These techniques will do little to improve that.
OTOH, you have control over how the code is vectorized (stripes or tiles), that could make a difference. And it might be worth exploring if it’s worthwhile to emulate doubles with smaller types. GPU’s are slow at 64 bit math.
But, as the author already suggested, first check if there is a need to optimize at all.