4 ms·
The most frequent application is machine learning. If my piece of code is implementing some function that I want to minimize, then taking the derivative (w.r.t.
by srl 6y ago
The most frequent application is machine learning. If my piece of code is implementing some function that I want to minimize, then taking the derivative (w.r.t. some vector of parameters) tells me in which direction I need to change the parameters in order to reduce the function: "gradient descent".
- dnautics 6y agoNote that the example here is forward differentiation and most machine learning these days is backward propagation
- oxinabox 6y agoNote that this applies backwards also. Of the 7 ADs demoed, only 1 was forward. The other 6 were reverse mode. All gave same result
- dnautics 6y agoOf course, sorry didn't mean to imply that AD was useful only in one direction. Haha I guess I was caught only reading 1/7 of the article!!
- rockinghigh 6y agoBackward propagation computes the gradient of the loss function with respect to the weights. A gradient is a vector or tensor of derivatives. You need to differentiate the loss function somehow to evaluate this vector.
- dnautics 6y agoThat's not strictly speaking true. You also calculate the gradient of the input. Otherwise, you wouldn't be able to backpropagate over more than one layer. It's just useful because you can do both easily, where with forward mode getting derivatives with respect to weights is considerably more expensive.
- jpollock 6y agoSo, we've got "reality", as represented by data. We've got a model, implemented in code. Since it's code can be differentiated - not sure how that works with branches, I guess that's the math. :) This is generated through a set of input parameters. We've got an error function, representing the difference between the model and reality. If we differentiate the error function, we can choose which set of parameter mutations are heading in the right direction to then generate a new model? We check each close point and find the max benefit? However, if everything is taking the parameters as input, is the derivative of the error function only generated once? Is it saying that the derivative of the error function is independent of the parameters, so it doesn't matter what the model is, they all have the same error function, and that error function can be found by generating a single model?
- eigenspace 6y agoDealing with branches is indeed and interesting problem. Many AD systems can't accomodate branches. I think most of the Julia ones do. The gist of it is that branches don't actually require that much fanciness, but they can introduce discontinuties, and AD systems will often do things like happily give you a finite derivative right at a discontinuity where a calculus student would normally complain.