3 ms·
That's not quite what they're talking about WRT zero human knowledge. The problem is that there's no intrinsic scoring system for Go, nothing specific to maxim
by tyler_larson 9y ago
That's not quite what they're talking about WRT zero human knowledge.
The problem is that there's no intrinsic scoring system for Go, nothing specific to maximize, so it's difficult to tell a computer whether a given outcome is "good" or "bad". So early versions of AlphaGo used a collection of human-played Go games to get an idea of what constitutes "good" and what is "bad", so it can then train its model to predict whether a move will make things better or worse.
This new system forgoes that step, and instead has the model play itself starting at random and looking for patterns that end up winning games. It's as if you gave the rules to the game of Go to a culture that's never heard of it before, and they evolved their own play style entirely in isolation.
Their result is a model that is better than the one that was developed with human influence, and that's the interesting bit.
- sobellian 9y agoI understand that the paper means that they didn't train it on expert input. The significance of the research is that this is a more general way to construct a game AI. The question I am posing is how far we have to go on that front.
- tyler_larson 9y agoSo... like this? https://www.theverge.com/2017/8/9/16117850/deepmind-blizzard-starcraft-ai-toolset-api https://www.theverge.com/2017/8/9/16117850/deepmind-blizzard...
- sobellian 9y agoYes. It'll be interesting to see if their starcraft project uses the same algorithms or not. Note that the link merely describes software that could be used for feature engineering. It doesn't describe what NN architecture or tree search algorithms deep mind is using.