3 ms·
This is a topic I love to study. The mathematical analysis is reasonable; the policy gradient is a classic approach; I love Sutton’s RL book on it: http://incom
by espadrine 2y ago
This is a topic I love to study. The mathematical analysis is reasonable; the policy gradient is a classic approach; I love Sutton’s RL book on it: http://incompleteideas.net/book/RLbook2020.pdf http://incompleteideas.net/book/RLbook2020.pdf
Even though nowadays many people rather use cross-entropy for training RL, which I believe leads to more stable training.
Some of the equations feel a bit imprecise. I prefer to use random variables (rather than the U term), so that uncertainty permeates every value. As a result, the epsilon-t formula might be a perfectible fit (it goes negative past a certain t, which is unrealistic).