3 ms·
It used to be on par on nv hardware. Then nvidia just stopped improving their OpenCL backend. Otherwise assuming no fancy features are used they are identical
by sharpneli 8y ago
It used to be on par on nv hardware. Then nvidia just stopped improving their OpenCL backend.
Otherwise assuming no fancy features are used they are identical in their programming model.
- TomVDB 8y agoAFAIK the part of identical programming models isn’t the case anymore since Volta, because Volta and Turing have now progression guarantees when you have intra-warp divergence due to the presence of unique program counters for each thread. Before that, intro-warp divergence combined with badly placed synchronization operations could result in hard hangs. I don't think this makes a different in terms of raw low level performance, but it might have an impact in terms of implementation algorithms that require synchronization?