4 ms·
It is already mainly training by playing against itself: https://googleblog.blogspot.se/2016/01/alphago-machine-learning-game-go.html https://googleblog.blogsp
by krig 11y ago
It is already mainly training by playing against itself:
https://googleblog.blogspot.se/2016/01/alphago-machine-learning-game-go.html https://googleblog.blogspot.se/2016/01/alphago-machine-learn...
> To do this, AlphaGo learned to discover new strategies for itself, by playing thousands of games between its neural networks, and adjusting the connections using a trial-and-error process known as reinforcement learning.
- aquadrop 11y agoIt's still based on human games. It plays itself but the way it plays was inherited from human. I wonder if there is some fundamental barrier to what you can reach with reinforcement depending on your base.
- relic 11y agoIt is based on human games until it can explore well enough to sufficiently break away from local optimums.
- aflinik 11y agoHaving it learn on human games was just a way of speeding up the initialization process before running reinforcement learning, it didn't limit the state tree that was being searched later on.