5 ms·
We do know that AlphaZero's chess playing style is different. AlphaZero doesn't care as much about material, while Stockfish and humans (presumably learning fro
by backpropaganda 8y ago
We do know that AlphaZero's chess playing style is different. AlphaZero doesn't care as much about material, while Stockfish and humans (presumably learning from Stockfish) do. This is because Stockfish uses a material heuristic to search for moves.
Consider rock-paper-scissor, and agent X which plays X every move. Agent rock beats agent scissor all the time, and agent scissor beats agent paper all the time. But we cannot conclude that agent rock will beat agent paper all the time. Chess and Go might have this intransitivity property. The fact that one game exists which is not transitive means the burden of proof that Chess and Go are transitive rests on the authors.
My own suspicion is that AlphaZero should be weaker against humans compared to AlphaGo Lee since the former is not trained on human games while the latter is.
- gliptic 8y agoBut that isn't true. AlphaGo Zero seemed to be considerably stronger against humans than AlphaGo Lee. Lee managed to win one game against AlphaGo Lee, but AlphaGo Zero won 60 games in a row against top players, as well as all games against Ke Jie.
- backpropaganda 8y agoI think you're wrong. The version that defeated Ke Jie and 60 masters online was AlphaGo Master, which afaik did train on human games.
- gliptic 8y agoYou're right, I misremembered how Master was trained because they already had developed Zero by then. I still don't think there's any reason to believe Zero would do worse against humans than the Lee version. Zero was much stronger than Master, which was initially trained on human games. If the consequent self-play learning caused any non-transitive relation to humans, it sure wasn't evident when Master played against humans. So why would it show up in the learning of Zero given how much stronger it is than Master. You would have a better case if there was a reason to believe humans played anywhere near optimally.
- backpropaganda 8y agoBecause Zero was not trained on human games and Master was? Do we have any model which was trained only from self-play that can beat humans? I don't think so.
- gliptic 8y agoYou seriously don't think Zero can beat humans? 6 stone advantage over AG Lee that was trained on human games? If AG Lee behaves like a superstrong human, then Zero handily beating AG Lee is clear evidence that Zero would beat humans. I don't understand how your intransitivity is supposed to work when DeepMind didn't find any when training in different ways. Are they supposed to run a tournament just to show that a massive improvement that crushes all previous versions can still beat puny humans?
- backpropaganda 8y agoYour intuition is strongly assuming transitivity: "A beats B handily and B beats C handily, therefore this is clear evidence that A beats C handily". I already gave an example where this is not true, and I don't know why you're still doubling down on the same wrong intuition. > DeepMind didn't find any when training in different ways Intransitivity may not occur within AG versions but could occur when compared to humans. You need very different types of strategies for intransitivity to occur. > Are they supposed to run a tournament just to show that a massive improvement that crushes all previous versions can still beat puny humans? I think so, yes. We do this with Chess. Why should Go be different? Deepmind is aware of this problem. See gwern's links if you are interested.
- darepublic 8y ago> AlphaZero doesn't care as much about material, while Stockfish and humans (presumably learning from Stockfish) do. It's the other way around no? Stockfish's preference for material is hand-coded by humans... because our best wisdom values material highly.
- qubex 8y agoI’m no chess expert but I am a game theorist. I got the feeling that AlphaZero cares about materiel only insofar as it occupies positions and projects power; in this sense it is further in the abstract-strategy realm than grandmasters and human-coded chess programs are. Human-style players are like quantum theory: they’re concerned with stuff. AlphaZero is like general relativity: it’s concerned with bending the backdrop so mere stuff becomes pliant and goes where it wants... of course you need stuff to bend the backdrop, but beyond that... I don’t know if I’m making much sense, cockeyed metaphors are my only means of expressing my impressions.
- yazr 8y agoSimilar sentiment has been expressed by others. In the (simpler and less studied) game of Shogi, AG made very startling moves, which increase the number of vulnerabilities (and opportunities). AG is able to better balance all these combinatoric paths. Perhaps as humans, we have limited combinatoric capacity, and hence are biased to conservative moves to reduce our search space (so we get lower expectations since we have to bound our variance?). Whether this is less-strategic or simply less-bound-by-human-limitation is up to discussion. end-rant EDIT: Would be interested to hear your views about larger DRL domains. Drop me a line if u r interested gaxasit@getnada.com temp email.
- porpoisely 8y agoThat's an excellent point. When humans are taught chess, we are initially taught the value or power of the pieces. The queen ( being the most powerful ), then rook ... all the way down to the pawn. Stockfish is programmed to know the value of the pieces. But Alphazero isn't taught anything but the rules. If it has any values for the pieces, it would be self generated though I doubt it generates any values for the pieces since it really doesn't need to and the value of pieces varies depending on the position. AlphaZero does seem to have a style which favors positional development over value. It seems to choose pattern over material. I wonder what would happen if one set of kids was taught the value of the pieces and another set weren't taught the value. Would one group be better than the other? Is knowing the value of the pieces a help or an impediment to developing your chess skill, especially at such an early age?
- newen 8y agoYou can look at piece value as a heuristic for the strength of your position. The problem with purely positional play without regard for material is that you could miss some tactic that forces exchanges and suddenly you're down three pawns and losing. As humans we miss tactics a lot, which is probably why we value material as much as positional advantages.