3 ms·
You’re confusing two different concepts. AlphaGo learns a distribution where given a game state it generates a move that maximises its internal probability of v
by tempusalaria 3y ago
You’re confusing two different concepts. AlphaGo learns a distribution where given a game state it generates a move that maximises its internal probability of victory. There is a second “distribution” namely that any terminated sequence of go moves has an objective result.
AlphaGo samples from the latter distribution in a guided way (as the space of all go games is computationally intractable). It uses its learned distribution to do that guided sampling and uses the objective outcomes of the known distribution to inform its own learned distribution.
One way to think about this in the context of language modelling. Suppose I want to build a language model that says the word “goal” at least once every 2000 tokens generated. I could then repeatedly generate from the model and objectively score whether it has generated that word or not in each occurrence (the analogy of the finished go game). I then can use this objective scoring function to compete models against each other and do the alpha go style training. You can see here how the new training data is sampled from a different distribution than just regular language.
- cubefox 3y agoYour "goal" example sounds like a more useless version of my theorem prover AI?