3 ms·
This is an overly restrictive view of RL. Yann's claim about the potential utility of RL, taken at face value, is clearly false. Backprop on a deterministic com
by psb217 11y ago
This is an overly restrictive view of RL. Yann's claim about the potential utility of RL, taken at face value, is clearly false. Backprop on a deterministic computation graph is equivalent to deterministic policy gradient in the same graph, where the reward is given by the value of the optimization objective.