4 ms·
Numba is great! Whenever you’re using CPU and have very simple parallelism patterns, it’s your best bet for speeding up numpy. But if you needed to do this on
by vladf 5y ago
Numba is great! Whenever you’re using CPU and have very simple parallelism patterns, it’s your best bet for speeding up numpy.
But if you needed to do this on a GPU or TPU, ideally with native and transparent SIMT, such as the case for the SO question inspiring the post (an unsupervised centroid-based loss for a deep learning setting), would you have to write a custom C++ kernel to do this?
Free SIMT may even make this worthwhile in the few-centroid setting.