7 ms·
A lot of the graphs in the paper seem to level out as they hit the level of the opponent. It makes me wonder to what extent AlphaGo Zero is merely optimizing to
by firebones 9y ago
A lot of the graphs in the paper seem to level out as they hit the level of the opponent. It makes me wonder to what extent AlphaGo Zero is merely optimizing to beat flaws in existing opponents' current implementations (even if "existing opponents" == all available opponents' data and algorithms today) rather than generalizable insights into the underlying game. Because wouldn't you expect that unless we are at the theoretical limit of perfect chess that a tabula rasa approach might exceed existing best practice significantly, especially with the massive computation advantage it has?
Not that there's anything wrong with that; AlphaGo Zero supposedly optimized for the "just enough" win rather than the crushing win. It doesn't even mean Stockfish is doomed--I suspect Stockfish could beat it in a future heads up match provided that Zero didn't have time to retrain, but that a retrained Zero (having the benefit of optimizing against a new Stockfish) would be able to supersede it once again.
- Houshalter 9y agoIt's not. It learns entirely through self play and never learns from playing it's opponent. Diminishing returns isn't unusual and happens in every domain. These AIs are probably playing close to the limit of what is possible, just not quite there yet.
- tlb 9y agoAre there popular games where the best human players are not near the limit of what is possible? Obviously you can construct one to be hard for humans (large 3SAT problems, or even big arithmetic problems), but I wonder if there is one that people enjoy.
- throw_away_777 9y agoHumans are nowhere near the limit of what is possible in chess, as evidenced by how much better computers are at the game.
- nandemo 9y agoPresumably tlb meant what is humanly possible...
- PeterisP 9y agoI'd assume that for pretty much any nontrivial game the best human players are nowhere near the limit of what's possible. Humans can play a perfect tic-tac-toe, but for everything in the realm of go, chess, poker, bridge, etc the theoretical ideal is far beyond currently best human players.
- gwern 9y ago> A lot of the graphs in the paper seem to level out as they hit the level of the opponent. DM is no longer investing much in the AG research program; Silver said the team has been disbanded already. If you look at the Go graph in this or the first AG0 paper, Zero was still getting better at Go when they shut it down, it hadn't converged. They just didn't want to tie up the TPUs. I don't think it's a coincidence that the graphs tend to stop after they reach superiority. (Also, as Houshalter says, one of the critical aspects is that this is pure self-play ie the NNs never play against the existing engines except for evaluation. So it's all independent from-scratch reinvention.)
- visarga 9y agoSeems like it flattens, but they only trained for a few hours. What would happen with 100x more training?
- jtolmar 9y agoELO ratings level out eventually for a given pool of opponents. If a player already wins every game against all available opponents, there's no evidence that can tell you if they suddenly got twice as good. If tracking improvements past the state of the art is important I think they'd have to freeze the algorithm every 400 ELO or so and rate the improved versions against the last snapshot. (Doesn't really apply to the stockfish case, but it does to the other two games.)