4 ms·
If you use autodiff and don't have an 1:1 relationship between function and argument indices you might have to do atomic sums or locking when computing the deri
by surban 9y ago
If you use autodiff and don't have an 1:1 relationship between function and argument indices you might have to do atomic sums or locking when computing the derivative because multiple elements of the function derivative correspond to one element of the argument derivative. On a CPU this might be okay but on CUDA GPUs this usually has a performance impact. Thus we transform the derivatives so that we have an explicit expression for each derivative element and thus can use one CUDA thread per derivative element.
- throwaway613834 9y agoOh wow, interesting. Thanks!