3 ms·
ADP = Reinforcement Learning. This got some attention lately when some people at google wrote a neural-network based RL system to play Atari games. https://ww
by idanoeman 11y ago
ADP = Reinforcement Learning. This got some attention lately when some people at google wrote a neural-network based RL system to play Atari games.
https://www.cs.toronto.edu/~vmnih/docs/dqn.pdf https://www.cs.toronto.edu/~vmnih/docs/dqn.pdf
- versteegen 11y ago(That was before Google acquired DeepMind, and probably a large part of the reason.) This uses a deep convolutional NN as an approximation function of the expected value. Interestingly there are only 3 hidden layers in this CNN, and it only gets fed the last 4 frames as state. Of course normally you would feed your approximator a direct encoding of the state. If you wanted to use a NN to evaluate nodes in gametree search you would have to invent some other way to train it. Would love to know whether there is research on that.