44 ms·
I think it's quite different. For GPUs these archiecture-specific optimizations / primitives make orders of magnitude difference (and also change from version t
by moab 4y ago
I think it's quite different. For GPUs these archiecture-specific optimizations / primitives make orders of magnitude difference (and also change from version to version). On the other hand, multicore code I wrote ~8 years ago in TBB or Cilk is still extremely fast compared to hand-optimized code customized for current CPU architectures.
To summarize my gripe, it's is that the abstractions on GPUs seem broken, and we still don't have a good model for how to write high performance GPU code that just works and keeps getting faster as GPUs get faster.