4 ms·
> to minimize the amount of computation IMO backprop is the most trivial implementation of differentiation in neural networks. Do you know an easier way to com
by om8 2y ago
> to minimize the amount of computation
IMO backprop is the most trivial implementation of differentiation in neural networks. Do you know an easier way to compute gradients with larger overhead? If so, please share it.
- QuadmasterXLII 2y agoMy first forays into making neural networks used replacement rules to modify an expression tree until all the “D” operators went away, but that takes exponential complexity in network depth if you aren’t careful. Finite differences is linear in number of parameters, as is differentiation by Dual Numbers
- eli_gottlieb 2y agoBackprop is the application of dynamic programming to the chain rule for total derivatives, which sounds trivial only in retrospect.
- Jensson 2y agoYou can do forward propagation. Humans typically finds forward easier than backwards.
- tripzilch 2y agosince you asked ... how about Monte Carlo with Gibbs sampling?