4 ms·
The error in backprop is a vector quantity, not a single scalar for each time step. In RL the goal is to optimize the overall sum of the reward over all time s
by jhartmann 11y ago
The error in backprop is a vector quantity, not a single scalar for each time step. In RL the goal is to optimize the overall sum of the reward over all time steps. Backprop attempts to minimize the magnitude of the error for a given loss function, by moving in the negative direction of the gradient of the function. Backprop just moves a lot more variables to a desired outcome, that is what LeCun is saying. The representational power of a single scalar 'score' doesn't have much ability to optimize such a large n-dimensional function in any efficient way.
- vonnik 11y agoMaybe I'm thinking about things wrong, but can't you have a scalar quantity to represent error for backprop at the end of a neural network (in a supervised regression problem for example)? That scalar error becomes a vector as it is assigned to various weights. I'm not trying to be obtuse, but it seems like the scalar reward in RL is also modifying countless variables in the Q functions of the state-action pairs that led to the final outcome/reward... Are those, by definition, smaller in number than the parameters of a neural network?