4 ms·
The gradient we want is the gradient with respect to the process which generated the dataset. The gradient we get is an estimate based on only a handful of samp
by LeegleechN 5y ago
The gradient we want is the gradient with respect to the process which generated the dataset. The gradient we get is an estimate based on only a handful of samples from that process at a time. The analogy holds up fine.
- eutectic 5y agoI would say it's more a commentary on the fact the the gradient is effectively based on L2 distance in parameter space, and so can be a bad/inefficent direction to move in even if you have access to the full gradient. Hence the motivation for momentum and second-order optimization.
- jmmcd 5y agoIf we were trying to describe stochastic gradient descent, this would be relevant, but we're talking about backprop (which does often use just a batch, but that's not inherent). And there is nothing about backprop that makes the kangaroo more blind than in any other form of gradient-based optimisation.