10 ms·
Show HN: AlphaZero Science paper
- RivieraKid 8y agoFinally! I've been waiting for this for a year. I love that it learned from scratch without human bias and how it plays in a much more captivating style than alpha-beta engines. Will you be answering questions?
- detaro 8y agoI don't see anything I can try out? https://news.ycombinator.com/showhn.html https://news.ycombinator.com/showhn.html
- gus_massa 8y agoFrom the profile, it looks that it may be one of the authors. In that case this can be more like a AMA than a ShowHN.
- sctb 8y agoMaybe a little borderline, but it seems like there are lots of resources to read and game data to play with, so we can probably spring for a “Show HN” in this case.
- stabbles 8y agoSo they finally released more games?! Really looking forward to Kingscrusher or Chessnetwork covering more of these 210 games on YouTube: https://deepmind.com/research/alphago/alphazero-resources/ https://deepmind.com/research/alphago/alphazero-resources/
- gpm 8y agoAgadmator covered one already: https://www.youtube.com/watch?v=ZHfumZVPjVA https://www.youtube.com/watch?v=ZHfumZVPjVA
- conistonwater 8y agoDaniel King is doing them now (https://www.youtube.com/watch?v=pFtY7gNRVRI https://www.youtube.com/watch?v=pFtY7gNRVRI), and so is Matthew Sadler for Chess24 (https://www.youtube.com/playlist?list=PLAwlxGCJB4NchyTBYik8FBbnzLXpCCO79 https://www.youtube.com/playlist?list=PLAwlxGCJB4NchyTBYik8F...).
- gamegoblin 8y agoIt's a shame they still played against 2016 Stockfish (Stockfish 8), when Stockfish 9 or Stockfish Dev were available (Stockfish 10 is out now, but only very recently, so I can understand why they didn't use it). Their results show that they are only just barely stronger than Stockfish 8, but Stockfish 9 and 10 are stronger than 8 as well. EDIT: Also meant to include a shout-out to http://www.lczero.org/ http://www.lczero.org/ which is an open source implementation of AlphaZero chess. Here is their forum post for this paper: https://groups.google.com/forum/#!topic/lczero/TfmaNHI99gk https://groups.google.com/forum/#!topic/lczero/TfmaNHI99gk SECOND EDIT: I was wrong! They did play against a newer SF than 8, specifically, SF at this commit: https://github.com/official-stockfish/Stockfish/commit/b508f9561cc2302c129efe8d60f201ff03ee72c8 https://github.com/official-stockfish/Stockfish/commit/b508f... , which was about 2 weeks before SF 9 was released, so maybe it is close in strength to SF 9.
- DannyBee 8y agoThe paper was in peer review for a year, so it couldn't have compared against those. AFAICT based on dates, they chose nearly the latest version they could have.
- thom 8y agoI'll never forget the last round of games with everyone gushing over AlphaZero's wonder human-like play as it reached a drawn position only for Stockfish to blunder a whole queen with nobody remarking on it. I played along while watching the Danny King video: https://twitter.com/DanielKingChess/status/1070755986636488704 https://twitter.com/DanielKingChess/status/10707559866364887... Recent Stockfishes recommend many of the quiet, strategic moves that he seems particularly enamoured with.
- plopz 8y agoIt said it was winning games with 1:10-100 time control. Why would you say thats only just barely stronger?
- gamegoblin 8y agoThe Elo graph in Figure 1 of the paper suggests it is maybe 30 Elo higher (eyeballing the graph, as they don't provide exact figures).
- thom 8y agoAnother set of games against an outdated Stockfish which appears to make moves that a recent Stockfish at any reasonably depth disagrees with. I've no doubt at all that AlphaZero has a much stronger evaluation algorithm than Stockfish, but I do wish they'd be a bit more transparent about its actual strength (although presumably they're selling access to it right now if you connect all the dots).
- RivieraKid 8y agoThe paper used latest AlphaZero and Stockfish at the time of writing. It's been in peer-review for almost a year.
- thom 8y agoWell, let me be more blunt: there is zero chance they're playing fair here.
- DannyBee 8y agoFWIW: Accusations of bad faith with no backup are not generally good commentary. As someone said, it was in peer review for a year. That means it could not have been compared against stockfish 9 or 10 - they were not released yet. As someone above points out, they used in-development stockfish versions as well (2 weeks before sf9 was released), and from what I can tell, they used the newest version they could have. If you have some good data that means you can make this statement, can you please cite it? Otherwise, can you please not make claims of bad faith?
- thom 8y agoThe strength of the language I used was motivated largely by the obviously bad-faith approach employed when they _first_ announced AlphaZero a year ago. They wanted to be able to say they'd created the strongest engine in the world, and they created an environment where it was hard to fail. I will admit the setup used in this paper seems more reasonable on that basis. That said, I've checked out a few old versions of Stockfish (including the exact commit they used) and analysed the games in Table S6 in the paper. Stockfish still spots multiple blunders in its own play. Obviously these things aren't entirely deterministic, but it seems unlikely there was time trouble. And again, just to be clear, there's little doubt in my mind that AlphaZero's evaluation of any given position is better than Stockfish's or anyone else's. I'm not even saying it can't reliably beat Stockfish. I just find it sad that the evidence of its overall strength continues to be wobbly.
- YetAnotherNick 8y agoWhy Show HN? It generally implies a single person or just a few people behind it.
- mcphage 8y ago> AlphaZero and AlphaGo Zero used a single machine with 4 first-generation TPUs and 44 CPU cores. A first generation TPU is roughly similar in inference speed to commodity hardware such as an NVIDIA Titan V GPU, although the architectures are not directly comparable. > The amount of training the network needs depends on the style and complexity of the game, taking approximately 9 hours for chess, 12 hours for shogi, and 13 days for Go. How much would that much computing power would cost on something like AWS? That's a lot of hardware, but if you're only renting it for 9 hours... the beefiest EC2+GPU instance Amazon has currently is p3.16xlarge, which has 8 Tesla V100 GPUs, and 64 (virtual) CPUs, for $25/hour on-demand. My understanding is that a V100 is slightly more powerful than a Titan V, so does that mean you could run the Chess training (at least the AlphaZero side) for $225? That seems impossible? EDIT: pacala below pointed out that the hardware listed was just for running AlphaZero against Stockfish, not for training it. Digging through the preprint itself, they say that for training they used: > During training only, 5,000 first-generation tensor processing units (TPUs) (19) were used to generate self-play games, and 16 second-generation TPUs were used to train the neural networks. So that would be... a lot more.
- pacala 8y agoThe most expensive part of training AlphaZero is creating the training dataset by self-playing tens of millions of games.
- chewxy 8y agoFirst game is cheap :P Subsequent games are expensive because scaling becomes an issue
- mcphage 8y agoAh! Okay, that's what I think I misunderstood. The 4 TPU + 44 GPU configuration was only for running AlphaZero against Stockfish, not for training it. Phew! That seemed unbelievable, I was hoping someone would see what I was missing. And the 9 hours was for training, but I don't think the article linked says on what.
- spac 8y ago
- schaefer 8y agoInufu, As a Go player, is there any way I could download and review the game records described in this paper?
- gwern 8y agohttps://deepmind.com/research/alphago/alphazero-resources/ https://deepmind.com/research/alphago/alphazero-resources/
- schaefer 8y agoI appreciate the link, but there doesn't appear to be a single Go game record at this link. Where as the paper describes a thousand (or more) games played between Alpha Go Zero and AlphaZero.
- gwern 8y agoMy point was that that is the official source for all the games they released so you can see for yourself. If they aren't there, they aren't there. You can download the earlier games (http://www.alphago-games.com/ http://www.alphago-games.com/) but apparently not these new ones.
- schaefer 8y ago> "If they aren't there, they aren't there" That's why I've specifically asked the author. Perhaps if Inufu sees there's interest, more game records could be released.
- gwern 8y agoIf you've been following the past DM releases, you'd know they pretty much never release additional material when asked and it's pointless asking a junior employee (who hasn't answered any questions to begin with). WYSIWYG.
- 8y ago
- mindgam3 8y agoTo OP or anyone else at DeepMind: can you comment on why you decided not to release all of the games? IMHO as a competitive scholastic chess player (former national U16 champion and top 3 world U10) and software engineer, it would significantly increase credibility of results. Not to mention would be fascinating to see the “ugly” games in addition to the ones handpicked by your team.
- ehsankia 8y agoThey've released over a 100 new ones here now. https://deepmind.com/research/alphago/alphazero-resources/ https://deepmind.com/research/alphago/alphazero-resources/
- mindgam3 8y agoThat’s a step, but I still find it weird to release some but not all of the games. I’m trying to come up with a logical reason other than they have something they don’t want people to see in the rest of the data, but so far I’m failing.
- faceplanted 8y agoIt doesn't necessarily come under "logical reasons", but the DeepMind team have pretty strict rules on data retention, chances are there's a debate in the company about it.
- achille 8y agoThere's a book coming out in a few weeks with more details. The authors were provided with all the data. What would you do with 1000 games that's not possible with 100?
- mindgam3 8y agoI would be able to accurately gauge AlphaZero’s true chess ability, strengths and weaknesses. Right now it’s impossible to do that given a curated selection of games. It would be like evaluating someone’s coding ability based on a few samples of their absolute finest work rather than the code they actually write day in and day out.
- nuguy 8y agoIt is my moral obligation to express to you the fact that AI, even this kind of AI, is a death sentence for humanity. The progress of automation will eventually meet and surpass the human mind. But even before it does, perhaps long before it does, it will cause massive economic disruption and unemployment. The more complete automation becomes, the less power humans will have, the less influence humans will have over the powerful entities that hold the keys to critical resources such as jobs. The economics of automation leave little doubt that the outcome will be bad for humans. I’m sorry I can’t explain it more effectively here. But I think it’s clear to anyone who thinks it through carefully. Please stop applying your intillenge to AI. Edit: substantive counter-arguments would be highly appreciated
- Tenoke 8y agoCan you imagine any ways in which that kind of AI can be used for good (or less bad), however unlikely? If yes, become an AI researcher/policy shaper or donate to groups that might be able to make progress towards the better outcomes.
- nuguy 8y agoThe individual use-cases don’t matter. What matters are the points of contact between AI as a concept and the fundamental economics of human life as we know it. AI will change the economics of life in very fundamental and very negative ways. Why does it matter or help to be active some branch of AI in particular? It won’t change this! The only solution to this problem is the banishment of AI. There is no other way to preserve life as we know it. AI might not provoke these changes within my lifetime. But people are very happy to protest and march for global warming even though it also will not end the world within our lifetime. There is a strange cognitive dissonance there. The logic is very similar: even if there is a small chance that it could end the world, better to err on the side of caution. The consequences of AI will be indescribably worse for humans than global warming, so why not exercise caution?
- feanaro 8y agoBecause the arguments for global warming eventually ending the world are so far much stronger. The endgame for AI is much less clear and therefore the final outcome much less certain.
- madisfun 8y agoIt would be nice if A0 participated in at least one public computer chess championship, Chess.com's CCC or TCEC. That's a level playing field and all games published. AlphaZero was a great concept and execution, but if we have to judge its relative strength, it should compete fairly. 4 TPUs (~ 4 Titan V) + 44 cores for AlphaZero vs only 44 cores for Stockfish pre-9 may or may not have put Stockfish at a disadvantage. BTW, current, presumably balanced, TCEC 14 configurations are: Non-GPU Server: CPUs: 2 x Intel Xeon E5 2699 v4 @ 2.8 GHz, Cores: 44 physical, RAM: 64 GB DDR4 ECC GPU Server: GPUs: 1 x 2080 ti + 1 x 2080, CPU: Quad Core i5 2600k, RAM: 16GB DDR3-2133 TCEC GPU server looks more modest than what A0 authors used to "beat" SF.
- ekelsen 8y agoThey also beat it with only 1/10th the time that stockfish got. That should more or less negate any advantage of the extra processing power.
- madisfun 8y agoOn paper, the results are amazing, but researchers are always biased to produce positive results. BTW, a match between StockFish 10 and LeelaChessZero, an open source implementation of the same idea, will be organized in a couple of days. From the LC0 blog: Lichess.org will host a match between the mighty Stockfish 10 and Leela. It will be a 6 games match with time control of 5'+2" with ChessNetwork commentary. Games will be played on 15th December at 17:00 UTC. Stockfish 10 will run on 64 cores 2.3GHz Xeon, while Leela will use the latest v19.1 Lc0 with 11248 network and will run on one GTX 1080 Ti + one RTX 2080 GPU.
- AndyNemmity 8y agoThe TCEC GPU overheated during the games, without anyone knowing until later. Then they underclocked it dramatically, as there was poor to no cooling in the datacenter they rented it from. That reason alone, if I were Deepmind, I would not be included in those competitions. It would be horrible press for them, that would involve a ton of human error out of their control.
- deleted 8y ago[deleted]
- antirez 8y agoHow it is possible that a so high profile article says a so doubtful statement? "Traditional chess engines – including the world computer chess champion Stockfish and IBM’s ground-breaking Deep Blue – rely on thousands of rules and heuristics handcrafted by strong human players that try to account for every eventuality in a game."
- have_faith 8y agoIt's not that far from the truth though is it? https://github.com/official-stockfish/Stockfish/blob/master/src/pawns.cpp https://github.com/official-stockfish/Stockfish/blob/master/...
- nmca 8y agoSo this is... Sort of true? Alpha/Beta pruning relies on heuristics and exploits their power with search. The heuristics that make stockfish so strong are numerous, handcrafted and try to account for many notions of material. In the current computer chess championships, there are Monte Carlo engines that use the search search strategy as AlphaZero with hand crafted heuristics, and they're doing ok. But they're not as strong as AZ, which learnt, by itself, what a good chess position looks like, starting from random play
- judofyr 8y agoLooking at the recent patches[1] for Stockfish it seems like a rather true statement? Most of the patches are related to how it evaluates a position; not how it prunes/searches the tree of possible moves. Here's one random test which showed improvement: https://github.com/Vizvezdenec/Stockfish/compare/5c2fbcd...51e1b4b https://github.com/Vizvezdenec/Stockfish/compare/5c2fbcd...5... Bitboard b1 = double_pawn_attacks_bb<Them>(nonPawnEnemies) & b & attackedBy[Us][PAWN]; score += make_score(90, 72) * popcount(b1); And look at all of the magic numbers and logic in evaluate.cpp: https://github.com/official-stockfish/Stockfish/blob/master/src/evaluate.cpp https://github.com/official-stockfish/Stockfish/blob/master/... [1]: http://tests.stockfishchess.org/tests http://tests.stockfishchess.org/tests
- antirez 8y ago
- Rampoina 8y agoWhat is Deepmind's interest in not releasing the source code and weights for the neural networks? I'm excited about their work but it seems that it would be much better for everyone if they just released their work openly.
- kevinwang 8y agoDunno if I just missed it in the paper, but is there an explanation for why alphazero is better at go than alphago zero?