3 ms·
AlphaZero plays games with (1) perfect information and (2) well-defined winning conditions. Neither of these hold for most human-learning scenarios. I can imag
by infinity0 9y ago
AlphaZero plays games with (1) perfect information and (2) well-defined winning conditions. Neither of these hold for most human-learning scenarios.
I can imagine that a healthy dose of probability theory (and probably more advanced stuff I don't know about[1]) might improve (1), but (2) is going to keep computer scientists and philosophers and ethicists arguing for quite a long time. :)
[1] get the joke, eh? eh? eh?
- abecedarius 9y agoThey essentially hold for math, which is a pretty big deal.
- infinity0 9y agoWhat are the winning conditions for math?
- ansible 9y agoA smaller proof using fewer axioms or other proofs than the current state-of-the-art. Discovering new and "interesting" proofs. Don't ask me to define "interesting" in this context.
- infinity0 9y agoIt's exactly my original point that these goals are not well-defined in early 21st century maths.
- seanwilson 9y agoIn the field of formal/machine proofs, nobody really cares about the length of the proofs because part of the point is the proofs are checked by the computer back to the basic axioms so you can trust the proofs are correct. Being able to discover long and ugly proofs to difficult theorems or coming up with new theorems would have endless applications.
- paulcole 9y ago> AlphaZero plays games with (1) perfect information I'm not sure why this matters? Everyone plays chess with perfect information. Both players see the entire board and all possibilities unlike, say, Scrabble or poker.
- theptip 9y agoI think GP meant that in the sense of "AlphaZero can only play games that have perfect information". It's a restriction of the algorithm, not a statement about how AlphaZero approaches the games it plays. This is why AlphaGo leveled up into AlphaZero playing Chess, and didn't learn to play Starcraft (yet).
- paulcole 9y agoah yeah, i gotcha. my bad
- irascible 9y agoI disagree that it plays with perfect information, in that it is playing against a different version of itself. It can "discover" that it is playing with perfect information, i.e. learn strategies that are more effective against a clone of itself, but it first has to develop some "sense" that this is the actual case.. and throwing in some random stock fish opponents or opponents that are clones of itself, but cloned at different stages of thier learning process, would remove this "perfect information" caveat. Unless perfect information only applies to the state of the board.. in which case, perhaps a "fog of war" visibility algorithm could push that boundary.. but then your talking about self learning with a blindfold type scenario.