Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mjaskowski
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
mjaskowski
11y ago
Actually I did both frame skipping and amxing out two frames as reported in the Natures letter https://storage.googleapis.com/deepmind-data/assets/papers/D...
2.
▲
by
mjaskowski
11y ago
Absolutely. Q-learning has this capabilities and a shallow neural network was used back in 1992 to play backgammon, which has a lot of stochasticity. See https://en.wikipedia.org/wiki/TD-Gammon
3.
▲
by
mjaskowski
11y ago
Yes, we try to approximate Q function with neural network. Which is basically an enhanced version of gradient-descent Sarsa. The main trick to notice is that you can't provide consecutive frames as mini-batches as these would be highl
4.
▲
by
mjaskowski
11y ago
Let me fix that. There are actually two papers: https://www.cs.toronto.edu/~vmnih/docs/dqn.pdf and a more recent and more detailed: https://storage.googleapis.com/deepmind-data/assets/pap
5.
▲
by
mjaskowski
11y ago
Note that Neural Network is just a very complex function. You usually think of Q as a function (S, A) -> (Expected accumulated future reward) which is equivalent to S -> A -> (Expected accumulated future reward) the Neural Network
6.
▲
by
mjaskowski
11y ago
I noticed that the video does not work in Safari. Interesting. I am not sure how this deepmind game was played. Note that a typical game with epsilon = 0.1 achieves results around 550 points. In the nature paper they were running the game w