4 ms·
Is this a library or something I can download and try training myself (on a small scale)? I'm not in a position to read the paper right now, so my apologies if
by Sukotto 9y ago
Is this a library or something I can download and try training myself (on a small scale)?
I'm not in a position to read the paper right now, so my apologies if that's covered in there. I want to ask just in case it's not, while this is still on the front page.
- gwern 9y agoNo. DM only occasionally releases software. Expert iteration is simple enough that someone can code it up on their own and there's already a few clones, so if anyone cares to train their own, it's doable, although it may take a while.
- chillee 9y ago"a while" is a bit of an understatement. Leela zero (the main alphago zero replication project) is a crowd sourced computation effort that's going to take a fairly long time to get anywhere. And from this paper: > "Training proceeded for 700,000 steps (mini-batches of size 4,096) starting from randomly initialised parameters, using 5,000 first-generation TPUs (15) to generate self-play games and 64 second-generation TPUs to train the neural networks."
- Houshalter 9y agoYou don't have to start from zero though. It's cool that it works with google scale resources. But it seems like it would be faster to initialize with a neural net first trained to mimic the moves of an existing chess or Go AI. And then improve it from there. >"Why is the net wired randomly?", asked Minsky. "I do not want it to have any preconceptions of how to play", Sussman said. Minsky then shut his eyes. "Why do you close your eyes?", Sussman asked his teacher. "So that the room will be empty." At that moment, Sussman was enlightened.
- nandemo 9y agoI'm pretty sure starting from zero is the point of the Leela-Zero. If they started from Stockfish, it wouldn't be a replication of AlphaZero.
- chillee 9y agoI don't think it's definitely true that will work well. AlphaZero did significantly better than the original versions of AlphaGo (which did learn from existing human games). However, even training those nets will still take a fairly intensive amount of computational resources. As for that koan, I'm not convinced it's very applicable here. My interpretation of the koan is that the entire setup (training process, structure, etc.) all encode domain knowledge. In this case, I think AlphaZero's domain knowledge is transferable enough that I don't think it's relevant.
- gcp 9y agoThe problem is that it isn't entirely clear whether this produces equal quality results. You might end up on a lower optimization plateau.