3 ms·
Isn’t it exactly what alphazero did? “AlphaZero was trained solely via self-play using 5,000 first-generation TPUs to generate the games and 64 second-generati
by d0mine 2y ago
Isn’t it exactly what alphazero did?
“AlphaZero was trained solely via self-play using 5,000 first-generation TPUs to generate the games and 64 second-generation TPUs to train the neural networks, all in parallel, with no access to opening books or endgame tables. After four hours of training, DeepMind estimated AlphaZero was playing chess at a higher Elo rating than Stockfish 8; after nine hours of training, the algorithm defeated Stockfish 8 in a time-controlled 100-game tournament (28 wins, 0 losses, and 72 draws).” [emphasis added] https://en.wikipedia.org/wiki/AlphaZero https://en.wikipedia.org/wiki/AlphaZero
- red75prime 2y agoI thought that it might be a rare chance to invoke the NFL theorem appropriately, but I guess I was wrong. The NFL talks about a uniform distribution of problems. A case that is probably never the case. At least for habitable universes. Nevertheless, the theorem basically states that there are games where AlphaZero will be beaten by another algorithm. Even if those games are nonsensical from our point of view.
- Xcelerate 2y ago> I thought that it might be a rare chance to invoke the NFL theorem appropriately, but I guess I was wrong Haha, I wouldn’t feel bad. It’s one of the most misunderstood theorems, and I don’t think I’ve ever seen it invoked correctly on a message board.
- tmtvl 2y agoI forget, was it Alpha or one of the others (Leela, Kata, FineArt,...) which had a weakness against... I wanna say the Micro Chinese (?), where it would consistently play the same suboptimal sequence that let players beat it easily if they took that path.
- voidmain 2y agoGames drawn from this uniform distribution can't even be implemented in our physical universe (you would need exponentially large lookup tables to store the rules). There is no chance of ever encountering any of them. Of course, there are "games" like "invert sha-512" that can be implemented in our world but are probably impractical to learn. But NFL has nothing to say about them; a game that simple has zero measure in a uniform distribution over problems.