3 ms·
We worked on a similar experiment two years ago (deep learning + reinforcement learning algorithm + some innovations, to learn to play Atari 2600 games). We obt
by mandor 13y ago
We worked on a similar experiment two years ago (deep learning + reinforcement learning algorithm + some innovations, to learn to play Atari 2600 games). We obtained similar scores in the games we tested but we did not submit any paper because we considered that the scores were not good enough. In particular, for Space Invaders, you can easily get 600 points by hiding behind a shelter while continuously firing, and never learn how to avoid the bullet.
So, I was not impressed by their results on Space Invaders.
Overall, we struggled to learn long-term strategies (finding pure reactive strategies is easy) and to learn to avoid bullets. They did too: "The games
Q*bert, Seaquest, Space Invaders, on which we are far from human performance, are more challenging because they require the network to find a strategy that extends over long time scales."
=> that's the real challenge...