3 ms·
It doesn't seem possible to give "no explicit rules", unless you count "making an illegal move equivalent to a loss" as giving no explicit rules. Which doesn't
by Tomminn 7y ago
It doesn't seem possible to give "no explicit rules", unless you count "making an illegal move equivalent to a loss" as giving no explicit rules. Which doesn't seem like anything but word laundering.
If you do any less than this, the net will be incentivized to make an illegal move for a win. In which case, yah, I'd guess that net would win a lot of chess games against other rule-bound nets.
- dang 7y agoWe changed the submitted title from "MuZero beats AlphaZero, with less training and no explicit rules: fully general" to that of the article.
- gwern 7y ago> It doesn't seem possible to give "no explicit rules", unless you count "making an illegal move equivalent to a loss" as giving no explicit rules. Which doesn't seem like anything but word laundering. The model is learning from game trajectories offline. Illegal moves will be assigned a very low probability because they do not appear in real games (whoever may be playing those games, whether in a simulation or in the real world). AlphaGo/AlphaZero did in fact mask out illegal moves in the MCTS tree search, but MuZero does not: > MuZero only masks legal actions at the root of the search tree where the environment can be queried, but does not perform any masking within the search tree.This is possible because the network rapidly learns not to predict actions that never occur in the trajectories it is trained on. And in ALE, all inputs are always legal, it's just that they may be useless and a waste of a 100ms turn. So for ALE it doesn't matter. Now, for Go/chess/shogi, they generate the training data for the supervised learning part by reusing MuZero for MCTS self-play. They don't mention how legal moves are handled there; they might be masking out like in AlphaZero. You could argue that this is, in some indirect way, 'explicit rules'. But since the MuZero is already learning to predict moves' value and which moves actually get taken, I see no reason that the MCTS self-play couldn't implement an instant-loss rule without any problem or slowing down training all that much, removing even that objection.
- sanxiyn 7y agoThey do mention "MuZero only masks legal actions at the root of the search tree... does not perform any masking within the search tree". So it's not an instant-loss rule.
- gwern 7y agoSearching within its internal game tree is different from what it actually does in the game-trajectory-generation phase. It could be entirely true that the search does not do anything at all to handle illegal moves, relying on the RNN being good enough to predict that illegal moves are never taken, and that the actual game-playing masks it out, or instant-loses, or something else entirely.
- pmontra 7y agoIllegal moves in Go do result in instant loss or are impossible. Examples 1. Taking a ko without playing elsewhere first: instant loss. 2. Playing on the top of another stone: in real world it's difficult to do because of the shape of the stones. They could make MuZero lose the game. 3. Playing when it's the opponent turn: instant loss. This is actually a way to resign: playing two stones together. Probably this is impossible to do for MuZero because the goal is to play one move. By the way, do those programs plan their next move even when the opponent is thinking, like humans do, or think only when it's their turn?