4 ms·
I did a great deal of reading on Q-learning around the time of the original AlphaGo, it looks like that was covered in a previous repo (RL-Adventures-1). This
by 2bitencryption 8y ago
I did a great deal of reading on Q-learning around the time of the original AlphaGo, it looks like that was covered in a previous repo (RL-Adventures-1).
This new one seems to not mention Q-learning, is that because all these examples are based implicitly on Q-learning, or are these totally new alternatives?
- blt 8y agoIn Deep Q Network (DQN), the Q network itself is the policy. You have one output for each possible action, and the the neural network estimates the Q value for each action in the current state. You act by selecting the action with the highest Q output. This doesn't work if the action space is continuous, e.g. motor torques for a robot. You might think you can fix this by making the action an input to the Q network and keeping only one output; then you could find the action with the highest output. But due to the nonlinearity in the neural network, this is an intractable nonconvex optimization problem. So instead, you train a neural network to output the action given the state. The algorithms are harder to understand, because Q learning is kind of like supervised learning but policy gradients really aren't. A lot of algorithms (A2C, DDPG, TRPO, etc.) still use one-output Q network (as described in the previous paragraph), but this is just a part of the learning algorithm * . Once training is done, you throw away this Q network. The learned behavior is entirely contained in the policy network. These methods are usually called policy gradient methods. This article covers policy gradient methods only. * it's possible to do "pure" policy gradients using only the empirical return, but the Q network helps reduce the variance of the gradient estimate and stabilize the learning.