9 ms·
Fast implementation of DeepMind's AlphaZero algorithm in Julia
- metalwhale 6y agoDisclaimer: I'm not the author. Just want to share this awesome project.
- mtgp1000 6y agoI don't know anything about Julia...how hard would this be to port to python or a c-style language? Edit: I was mainly asking because I was curious about the relative expressiveness Julia...
- instance 6y agoThere are already implementations out there in Python. [1] The point of that project is to be a very fast alternative to those implementations while being more accessible than a C++ implementation. [1] https://github.com/suragnair/alpha-zero-general https://github.com/suragnair/alpha-zero-general
- ViralBShah 6y agoI was going through this project over the weekend. And while I can't recall where exactly in the docs I read this, I am quite sure the author mentioned that there are various python projects but they are quite slow. Other implementations such as leela chess zero have a lot of C++ and are difficult to follow. In fact, one of the things we want to do is maximize the performance of the Julia implementation. We hope to co-develop the compiler and ML stack to address these issues as they come up.
- newswasboring 6y agoI don't know if you saw it here, but a similar point is made in the readme section "Why should I care about this implementation".[1] https://github.com/jonathan-laurent/AlphaZero.jl#why-should-i-care-about-this-implementation https://github.com/jonathan-laurent/AlphaZero.jl#why-should-...
- doublesCs 6y agoTruly truly thank you for your work <3
- newswasboring 6y agoNot sure you are thanking jonath_laurent (original author of package of discussion) or ViralBShah (co-creator of Julia). But I concur on both accounts :D
- doublesCs 6y agoViral. I thanked Jonath in a different post :-P When I see projects like this (I mean especially Julia, but also people sharing their work on packages like this) I feel very fortunate that elements of the free software movement are still alive.
- adamnemecek 6y agoWhy would you? Julia shines in this use case.
- cgreerrun 6y agoI've been working on a Python implementation that uses Gradient Boosted Decision Trees (LightGBM/Treelite) instead of using a neural network for the value/policy models: https://github.com/cgreer/alpha-zero-boosted https://github.com/cgreer/alpha-zero-boosted It's mostly to understand how AlphaZero&Friends work. I'm also curious about how well a GBDT could do, and if there are self-play techniques that can accelerate training. The nice thing about a GBDT is that, unlike when using a NN, you can do thousands of value/policy lookups per second on a single core. So it should be cheaper to scale self-play and run a lot of self-play experiments (assuming the self-play learnings when using the GBDT model transfer to when you use the more-powerful NN in these environments). If you're curious about accelerating self-play training, check out David Wu's work (https://arxiv.org/pdf/1902.10565.pdf https://arxiv.org/pdf/1902.10565.pdf). He's the creator of KataGo. I implemented his "Playout Cap Randomization" technique in my implementation above and, sure enough, it's much more efficient: https://imgur.com/a/epaKtDY https://imgur.com/a/epaKtDY. It seems like it's still early days in terms of how efficient self-play training is.
- jorgemf 6y agohow good is your AI so far?
- cgreerrun 6y agoI'm trying to answer that right now, actually. For connect 4, once it's trained a bit it seems to do really well. At 800 MCTS playouts (<.4s), it goes from "I can beat it and sometimes we draw" before training to "I pray for a draw, it almost always beats me" after ~4 hours of training. Connect 4 is a solved game, so it should be possible to sample the space of the trillions of (position, who should win?, what are the best move(s)?) tuples and compare the answers to your value/policy models to get some kind of objective error. I haven't had time to do that, but having that benchmark is nice to have so you don't have to do a "ladder tournament" against some reference bot(s) like you do for Go where you don't know what ideal play is. After training it for 10 hours on Quoridor (using my personal laptop), it still can't beat me, but it doesn't seem anywhere close to plateauing. It goes from the agents aimlessly wandering around the board looking for victory row and randomly placing walls, to putting walls that thwart the opponent and navigating to the victory row. I decided to implement PCR and try out some self-play techniques on Connect Four before I give it another go for Quoridor; a few days of self-play improvements can speedup training 10x. That's where I'm at now... Once I test a few strategies I was thinking of firing up a c5a24x, 96-core box on AWS and giving it another go. It's ~1-2$/hr at the spot price so I can probably do a lot of damage for 50$ or so.
- jonath_laurent 6y agoAuthor here: I am happy to answer any question you may have about AlphaZero.jl. :-)
- MaxBarraclough 6y agoHi, thanks for this great project. Connect Four was used as a demonstration. I presume this is because it's much easier/cheaper to train a Connect Four AI, compared to Go?
- jonath_laurent 6y agoYes. Go 19x19 would be completely intractable on a single machine (one comment is citing a $25 million cost estimate in computing power to train AlphaGo Zero). A more reasonable target would be Go 9x9 but even this would be an extreme challenge on a single machine. There is an Oracle blog article series about training a close-to-perfect Connect Four player using AlphaZero. Even here, they had to rely on multiple GPUs. You have to keep in mind that AlphaZero is an extremely sample-inefficient learning technique, even for simple problems. Rather, the strengths of this algorithm is that 1) it is pretty generic and 2) it can leverage huge amounts of computation.
- doublesCs 6y agoNo questions. Just wanted to thank you for sharing. People like you make the world better one tiny bit at a time.
- jonath_laurent 6y agoThanks for your kind message.
- patagurbon 6y agoDo you have any thoughts about multi GPU training? I haven't seen many options for Flux previously, but didn't dig very much.
- 6y ago
- tbenst 6y agoFirst of all this is very cool. Dunno if author is on here, but I’m curious why both Flux and Knet are used rather than just one of them (Flux seems the most Julianic?). Also, is this really faster than PyTorch/TF? Last time I benchmarked Flux for non-trivial networks, the speed was quite good with small models but memory usage was ~5x higher than pytorch, and I couldn’t fit my models on the GPU for flux. For large models, I had to compromise on batch size in Julia, although maybe with Zygote.jl the memory issues have been resolved?
- jonath_laurent 6y agoAuthor here. AlphaZero.jl supports both Flux and Knet indeed and users can choose whatever framework they want to use. As far as I understand, Flux and Knet have different strengths. I think Knet is a bit more stable and mature for large-scale Deep Learning, but Flux shines for "scientific-ML" usecases where low AD overhead is crucial.
- jonath_laurent 6y agoI suspect FLux/Knet are still slightly slower and less memory efficient than PyTorch/TF, although things are moving very fast here! This is not relevant in understanding AlphaZero.jl speed though. The reason it is much faster than Python implementations is because tree search is also a bottleneck, and Julia shines here!
- tbenst 6y agoAh, I hadn’t appreciated this. Thanks for making & sharing your code!
- ViralBShah 6y agoWhile some may be addressed and others are being addressed, what would really help us if people file issues when they don't find performance to be adequate. If you still have the code handy, please do open some issues.
- tbenst 6y ago
- tromp 6y agoThe implementation includes Connect Four as an example application. While the standard board size of 7x6 is indeed solved, as they note, and in fact all sizes up to 8x8 are [1], they could have picked 9x8 or 9x9 which are currently unsolved. The latter is the new standard size on Little Golem which upgraded from 8x8 when that was solved. [1] https://tromp.github.io/c4/c4.html https://tromp.github.io/c4/c4.html [2] http://www.littlegolem.net/jsp/games/gamedetail.jsp http://www.littlegolem.net/jsp/games/gamedetail.jsp? gtid=fir [3] http://www.littlegolem.net/jsp/forum/topic2.jsp?forum=80&topic=77 http://www.littlegolem.net/jsp/forum/topic2.jsp?forum=80&top...
- jonath_laurent 6y agoI completely agree with you. Let me just add two remarks. First, although picking 9x9 boards makes connect-four intractable for bruteforce search indeed, I would be suprised if it made it much more difficult for AlphaZero, which relies on the generalization capabilities of the network anyway. Second, using a solved game for the tutorial is a feature, not a bug. This allows precise benchmarking of the resulting agent as a ground truth is known.
- dnautics 6y agoThat's really cool and I didn't think of that. I just wanted clarification: that means you train the agent without the deterministic solution and your "validation/test" (I'm not sure what those phases are called in unsupervised learning) sets are done without the deterministic solution.
- jonath_laurent 6y agoYes, the agent is trained without access to the deterministic solution.
- tromp 6y agoI did not see an evaluation of how close to perfection the agent becomes. Did you compute any sort of error rate (by finding moves that turn a won position into a non-won one or a drawn position into a lost one) ? And how this error rate drops over time as learning advances? That would indeed be very interesting to see.
- FiberBundle 6y agoDoes anybody know how long it would take to train an alphazero go version using one gpu? In [1] they claim that it took 13 hours until the model was able to beat the original alphago version, but they don't state what hardware they used. [1] https://deepmind.com/blog/article/alphazero-shedding-new-light-grand-games-chess-shogi-and-go https://deepmind.com/blog/article/alphazero-shedding-new-lig...
- arijun 6y agoI can’t find it now but iirc there was a blog post on HN about a month ago that estimated their training costs at $25 million, using many TPU pods.
- cgreerrun 6y agoHere was the guestimation: https://www.yuzeh.com/data/agz-cost.html https://www.yuzeh.com/data/agz-cost.html
- newswasboring 6y agoFrom an offline chat with the original author, The ELF OpenGo paper[1], which is an open implementation of AlphaGo Zero developed by Facebook AI: "First, we train a superhuman model for ELF OpenGo. Af-ter running our AlphaZero-style training software on 2,000GPUs for 9 days, our 20-block model has achieved super-human performance that is arguably comparable to the 20-block models described in Silver et al. (2017) and Silveret al. (2018)." [1]: https://arxiv.org/pdf/1902.04522.pdf https://arxiv.org/pdf/1902.04522.pdf
- jonath_laurent 6y agoI agree with the quoted numbers. As I mentioned in another comment, you have to keep in mind that AlphaZero is an extremely sample-inefficient learning technique, even for simple problems. However, it has two major strengths: 1) it is pretty generic and 2) it can leverage huge amounts of computing power.
- klipt 6y ago
- likeaj6 6y agoThis is awesome! I worked on a similar project in the past for the game Hex Did a writeup here about it: https://notes.jasonljin.com/projects/2018/05/20/Training-AlphaZero-To-Play-Hex.html https://notes.jasonljin.com/projects/2018/05/20/Training-Alp... https://github.com/likeaj6/alphazero-hex https://github.com/likeaj6/alphazero-hex
- jonath_laurent 6y agoActually, I found your blog article when I was reading about AlphaZero and I found it useful!
- mahgoskska3660 6y agohttps://apps.apple.com/us/app/hack-for-hacker-news-developer/id1464477788 https://apps.apple.com/us/app/hack-for-hacker-news-developer...