3 ms·
I think it would just smoke you from the outset. As far as I know it doesn't have a structured intelligence it can scale back - it would make the optimal move e
by devindotcom 6y ago
I think it would just smoke you from the outset. As far as I know it doesn't have a structured intelligence it can scale back - it would make the optimal move every time, destroying you like it destroyed top-tier players.
I tried learning Go a little while back but hit a wall. Was thinking about trying this more gamified option:
https://www.wolfeystudios.com/TheConquestOfGo.html https://www.wolfeystudios.com/TheConquestOfGo.html
- plants 6y agoMy rudimentary understanding of most reinforcement learning systems is that there is an "probability of optimality" associated with each action. Wouldn't there be a way to make the AI take the Nth optimal move or vary the degree of optimality with each move?
- kadoban 6y agoFor these you are able to pick out worse moves than the optimal from them, but that's not actually the same thing as "play like a beginner, okay now play like an intermediate". These things are still openish questions and if nothing else there's a lot of room for improvement in tools to help you learn and review games.
- _hark 6y agoYou can make KataGo play moves that keep the score roughly even since it has a trained score head, e.g. kataJigo [1]. This will keep the game even to your level, a nice way to train. [1] https://github.com/sanderland/katrain#ais https://github.com/sanderland/katrain#ais
- kadoban 6y agoI'm not sure that's a good idea at all for training, though it is a really neat trick. For training you really want your good moves to be rewarded and your bad moves pointed out, but if the AI just plays up or down to match what you do instead, there's no signal getting back to you on how you're doing.
- Imnimo 6y agoIt turns out that taking a superhuman player and making it play like a weak human is surprisingly tricky. It's not so hard to make a weak player - you just take suboptimal moves instead of the best moves. But often these suboptimal moves are bizarre. A weak human chess player will lose their pieces as they fall prey to forks and skewers and so on - tricks that are hard to see coming for new player. But they will rarely actively throw a piece away by moving it into danger. Even a novice human is pretty decent at looking one move ahead. But to a chess engine, or an MuZero agent, a move that loses the queen immediately and a move that leads to a sequence that loses the queen in five turns are basically equal. And so an artificially-weak MuZero agent, or an artificially-weak Stockfish agent will tend to make 'mistakes' that not even a weak human would make. This makes them a little difficult to learn from. There does exist research on how to make a human-like weak player: https://arxiv.org/abs/2006.01855 https://arxiv.org/abs/2006.01855 The basic idea is to look at weak human games and try to predict when a mistake will be made. But I don't know if there's any approach that can do that without access to a corpus of human errors.