3 ms·
Deep dive tutorial for learning in a forward pass [1] [1] https://amassivek.github.io/sigprop https://amassivek.github.io/sigprop
by maurits 4y ago
Deep dive tutorial for learning in a forward pass [1]
[1] https://amassivek.github.io/sigprop https://amassivek.github.io/sigprop
- cochne 4y ago> There are many choices for a loss L (e.g. gradient, Hebbian) and optimizer (e.g. SGD, Momentum, ADAM). The output(), y, is detailed in step 4 below. I don't get it, don't all of those optimizers work via backprop?
- mkaic 4y agoThe optimizers take parameters and their gradients as inputs and apply update rules to them, but the gradients you supply can come from anywhere. Backdrop is the most common way to assign gradients to parameters, but other methods can work too—as long as the optimizer is getting both parameters and gradients, it doesn't care where they come from.