3 ms·
That got solved using Deep-Q learning. I think David Silver did loads of work on it? Basically, instead of computing every state value like in normal Q learnin
by silveraxe93 3y ago
That got solved using Deep-Q learning. I think David Silver did loads of work on it?
Basically, instead of computing every state value like in normal Q learning. You use a Neural Network to estimate the best value.
- visarga 3y agoIf I remember correctly DQLearning is using discrete action space, limited number of possible actions which makes it possible to simply take the argmax_i(qvalue(state, action_i))