4 ms·
This is a great post. However, it didn't touch on the two main problems of AD: - Using Dual Numbers requires that all functions that you call into accept temp
by Fede_V 10y ago
This is a great post. However, it didn't touch on the two main problems of AD:
- Using Dual Numbers requires that all functions that you call into accept templated parameters. If you want to use GSL, BLAS, or any other mature math library, you are probably out of luck.
- Even if you are willing to port the code and modify the functions to accept templated parameters, very highly optimized math libraries make assumption not just about the behaviour of numbers (their API, defined by how they behave under addition/subtraction, etc) but also about their ABI. For example, a well tuned LAPACK like OpenBlas or MKL has very well tuned loop sizes to optimize cache behaviour assuming that floats are of a particular size.
- imurray 10y agoIf the function is analytic and supports complex numbers (usual in Matlab/Numpy/Fortran/...), but you're only interested in the real part, a quick hack is to abuse complex numbers and compute imag(f(x+i*epsilon))/epsilon. That one-liner can often give the right answer to machine precision for small epsilon. http://blogs.mathworks.com/cleve/2013/10/14/complex-step-differentiation/ http://blogs.mathworks.com/cleve/2013/10/14/complex-step-dif... For matrix-based code, a good note on derivative propagation is https://people.maths.ox.ac.uk/gilesm/files/NA-08-01.pdf https://people.maths.ox.ac.uk/gilesm/files/NA-08-01.pdf — the symbols with dots on top are forward-propagated derivatives like dual numbers. The symbols with bars on top are back-propagated derivatives, which is what you want to compute lots of derivatives at once. "Machine learning" libraries like Theano, TensorFlow, and Autograd will compute the reverse mode operation for many linear algebra expressions automatically, and use reasonable libraries like BLAS+LAPACK or Eigen under the hood.