4 ms·
Comparing AlphaZero to AlphaGo Lee seems problematic. I don't think transitivity holds in this case, i.e. AlphaZero can beat AlphaGo Lee almost surely, and Alph
by backpropaganda 8y ago
Comparing AlphaZero to AlphaGo Lee seems problematic. I don't think transitivity holds in this case, i.e. AlphaZero can beat AlphaGo Lee almost surely, and AlphaGo Lee can beat Lee Sedol almost surely, but it could be possible that AlphaZero is not able to beat Lee Sedol at all. This is because the state spaces reached in computer-computer games are probably very different from the state space reached in human-computer games. I could be wrong, but at the very least this should be discussed by Deepmind.
- Wowfunhappy 8y ago> This is because the state spaces reached in computer-computer games are probably very different from the state space reached in human-computer games. Why would this be true? If we were talking about ai-generated music that would be one thing, but it isn't intuitively obvious to me why a computer would play Chess or Go all that differently from a human.
- backpropaganda 8y agoWe do know that AlphaZero's chess playing style is different. AlphaZero doesn't care as much about material, while Stockfish and humans (presumably learning from Stockfish) do. This is because Stockfish uses a material heuristic to search for moves. Consider rock-paper-scissor, and agent X which plays X every move. Agent rock beats agent scissor all the time, and agent scissor beats agent paper all the time. But we cannot conclude that agent rock will beat agent paper all the time. Chess and Go might have this intransitivity property. The fact that one game exists which is not transitive means the burden of proof that Chess and Go are transitive rests on the authors. My own suspicion is that AlphaZero should be weaker against humans compared to AlphaGo Lee since the former is not trained on human games while the latter is.
- gliptic 8y agoBut that isn't true. AlphaGo Zero seemed to be considerably stronger against humans than AlphaGo Lee. Lee managed to win one game against AlphaGo Lee, but AlphaGo Zero won 60 games in a row against top players, as well as all games against Ke Jie.
- backpropaganda 8y agoI think you're wrong. The version that defeated Ke Jie and 60 masters online was AlphaGo Master, which afaik did train on human games.
- gliptic 8y agoYou're right, I misremembered how Master was trained because they already had developed Zero by then. I still don't think there's any reason to believe Zero would do worse against humans than the Lee version. Zero was much stronger than Master, which was initially trained on human games. If the consequent self-play learning caused any non-transitive relation to humans, it sure wasn't evident when Master played against humans. So why would it show up in the learning of Zero given how much stronger it is than Master. You would have a better case if there was a reason to believe humans played anywhere near optimally.
- backpropaganda 8y agoBecause Zero was not trained on human games and Master was? Do we have any model which was trained only from self-play that can beat humans? I don't think so.
- gliptic 8y agoYou seriously don't think Zero can beat humans? 6 stone advantage over AG Lee that was trained on human games? If AG Lee behaves like a superstrong human, then Zero handily beating AG Lee is clear evidence that Zero would beat humans. I don't understand how your intransitivity is supposed to work when DeepMind didn't find any when training in different ways. Are they supposed to run a tournament just to show that a massive improvement that crushes all previous versions can still beat puny humans?
- backpropaganda 8y agoYour intuition is strongly assuming transitivity: "A beats B handily and B beats C handily, therefore this is clear evidence that A beats C handily". I already gave an example where this is not true, and I don't know why you're still doubling down on the same wrong intuition. > DeepMind didn't find any when training in different ways Intransitivity may not occur within AG versions but could occur when compared to humans. You need very different types of strategies for intransitivity to occur. > Are they supposed to run a tournament just to show that a massive improvement that crushes all previous versions can still beat puny humans? I think so, yes. We do this with Chess. Why should Go be different? Deepmind is aware of this problem. See gwern's links if you are interested.
- darepublic 8y ago> AlphaZero doesn't care as much about material, while Stockfish and humans (presumably learning from Stockfish) do. It's the other way around no? Stockfish's preference for material is hand-coded by humans... because our best wisdom values material highly.
- qubex 8y agoI’m no chess expert but I am a game theorist. I got the feeling that AlphaZero cares about materiel only insofar as it occupies positions and projects power; in this sense it is further in the abstract-strategy realm than grandmasters and human-coded chess programs are. Human-style players are like quantum theory: they’re concerned with stuff. AlphaZero is like general relativity: it’s concerned with bending the backdrop so mere stuff becomes pliant and goes where it wants... of course you need stuff to bend the backdrop, but beyond that... I don’t know if I’m making much sense, cockeyed metaphors are my only means of expressing my impressions.
- yazr 8y agoSimilar sentiment has been expressed by others. In the (simpler and less studied) game of Shogi, AG made very startling moves, which increase the number of vulnerabilities (and opportunities). AG is able to better balance all these combinatoric paths. Perhaps as humans, we have limited combinatoric capacity, and hence are biased to conservative moves to reduce our search space (so we get lower expectations since we have to bound our variance?). Whether this is less-strategic or simply less-bound-by-human-limitation is up to discussion. end-rant EDIT: Would be interested to hear your views about larger DRL domains. Drop me a line if u r interested gaxasit@getnada.com temp email.
- porpoisely 8y agoThat's an excellent point. When humans are taught chess, we are initially taught the value or power of the pieces. The queen ( being the most powerful ), then rook ... all the way down to the pawn. Stockfish is programmed to know the value of the pieces. But Alphazero isn't taught anything but the rules. If it has any values for the pieces, it would be self generated though I doubt it generates any values for the pieces since it really doesn't need to and the value of pieces varies depending on the position. AlphaZero does seem to have a style which favors positional development over value. It seems to choose pattern over material. I wonder what would happen if one set of kids was taught the value of the pieces and another set weren't taught the value. Would one group be better than the other? Is knowing the value of the pieces a help or an impediment to developing your chess skill, especially at such an early age?
- newen 8y agoYou can look at piece value as a heuristic for the strength of your position. The problem with purely positional play without regard for material is that you could miss some tactic that forces exchanges and suddenly you're down three pawns and losing. As humans we miss tactics a lot, which is probably why we value material as much as positional advantages.
- gwern 8y agohttps://arxiv.org/abs/1806.02643 https://arxiv.org/abs/1806.02643 and https://arxiv.org/abs/1803.06376 https://arxiv.org/abs/1803.06376 (as does the choice to drop the checkpoints & historical self-plays in general from the training) indicate that AlphaGo versions, at least up until then, tend to be transitive: > What is worthwhile to observe from the AlphaGo dataset, and illustrated as a series in Figures 3 and 4, is that there is clearly an incremental increase in the strength of the AlphaGo algorithm going from version αr to αrvp, building on previous strengths, without any intransitive behaviour occurring, when only considering a strategy space formed by the AlphaGo versions.
- backpropaganda 8y agoThanks a lot for the links. They look quite interesting. It does seem that Deepmind is aware of this, and are working on evaluating this. Transitivity might be true within AlphaGo versions, but that doesn't give me any confidence that it would also hold when a human is in the equation. If a group of policies more or less occupy the same state space, they are likely to be transitive, but if they occupy disjoint state spaces, I don't think we can be sure of transitivity.