10 ms·
The neural network of the Stockfish chess engine
- knuthsat 6y agoThis is very nice. Reminds me a lot of tricks used for simple linear models. And it seems to work, given that Leela Chess Zero is losing with quite a gap. Most of the times one would learn a model by just changing a single feature and then doing the whole sum made no sense. A good example is learning a sequence of decisions, where each decision might have a cost associated to it, you can then say that the current decision depends on a previous one and vary the previous one to learn to recover from errors. If previous decision was bad, then you'd still like to make the best decision for current state.). So even if your training data does not have this error-recovery example, you can just iterate through all previous decisions and make the model learn to recover. An optimization in that case would be to just not redo the whole sum (for computing the decision function of a linear model).
- kohlerm 6y ago"Leela Chess Zero is losing with quite a gap" Check https://tcec-chess.com/ https://tcec-chess.com/ ATM Leela is in front again. That being said it is not clear which approach is better e.g. a very smart but relatively slow evaluation function (Leela) or SFs approach to use a relatively dumb eval function with a more sophisticated search. It is pretty clear that Leelas search can be improved (Check https://github.com/dje-dev/Ceres https://github.com/dje-dev/Ceres for example)
- kohlerm 6y agoThat being said, assuming you can fully saturate multiple GPUs then the SF approach has the disadvantage that it cannot use the performance improvements for GPUs/TPUs which still are growing fast, whereas CPU performance only grows slowly these days.
- confuseshrink 6y agoInteresting point. Nvidia have been improving the int performance for quantized inference on their GPUs a lot. It might be a lot of work but could it be possible to scale up this NNUE approach to the point where it would be worthwhile to run on a GPU? For single-input "batches" (seems like this is what's being used now?) it might never be worthwhile but perhaps if multiple positions could be searched in parallel and the NN evaluation batched this might start to look tempting? Not sure what the effect of running PVS with multiple parallel search threads is. Presumably the payoff of searching with less information means you reach the performance ceiling quite a lot quicker than MCTS-like searches as the pruning is a lot more sensitive to having up-to-date information about the principal variation. Disclaimer: My understanding of PVS is very limited.
- kohlerm 6y agoSure if someone can come up with an approach to run an NNUE (efficiently updatable) network on GPUs that might really be another breakthrough. But at a first glance it looks to me that this could be very difficult. Because AFAIK the SF search is relatively complicated. Even for Leela implementing an efficient batched search on multiple GPUs seems to be difficult (some improvements coming with Ceres). And Leela is using a much simple MCTS search. That doesn't that Leela's search could not be improved. It does not give higher priority necessarily for forced sequence of moves (at least not explicitly) or high risk moves. Which is IMHO why sometimes she does not see relatively simple tactics.
- fho 6y agoI guess the simplest approach to port NNUE to GPUs would be to run a complete instance per GPU thread (ie concurrent, not parallel evaluation).
- stabbles 6y agoThe Stockfish mentality is: not everybody owns a GPU and Stockfish should be available for everybody and perform well [1]. So they went for a CPU micro-arch optimized neural net, which is great. Maybe this is ultimately not the best for tournaments. What I would find interesting is if they could give engines an energy budget instead of a time limit. Maybe that makes CPU vs GPU games more interesting & fair. [1] https://github.com/official-stockfish/Stockfish/issues/2823 https://github.com/official-stockfish/Stockfish/issues/2823
- knuthsat 6y agoLooks like the tournament is still ongoing. I was referring to the last few. https://en.wikipedia.org/wiki/TCEC_Season_18 https://en.wikipedia.org/wiki/TCEC_Season_18 https://en.wikipedia.org/wiki/TCEC_Season_19 https://en.wikipedia.org/wiki/TCEC_Season_19
- Fragoel2 6y agoI don't know a lot about chess but I have one question: isn't chess a solved game (in the sense that given a board state we always can compute the right move)? why use a neural network that can introduce, a very small percentage of the time, mistakes? I guess it is for performance reasons?
- elcomet 6y agoChess is definitely not a solved game. The end game is solved though, when you have only 7 pieces left on the board. See https://en.wikipedia.org/wiki/Endgame_tablebase#Computer_chess https://en.wikipedia.org/wiki/Endgame_tablebase#Computer_che...
- mvanaltvorst 6y agoYou can imagine playing chess as a tree. You start with a game state, and every possible move is an edge to a new game state. Unfortunately, the amount of game states grows approximately exponentially (e.g. for every game state there are approximately 40 moves to play, thus every extra move multiplies #(game states) by 40, approximately). This neural engine is a trick so that Stockfish does not have to simulate all 40 moves every game states. The neural net outputs the moves that are most likely to be strong moves, and Stockfish will only consider those moves. Of course, this is a very basic explanation and there are many more optimisations Stockfish uses, though the ultimate goal of almost every optimisation is to reduce the amount of simulation Stockfish has to do.
- V-2 6y ago"This neural engine is a trick so that Stockfish does not have to simulate all 40 moves every game states." This optimization always done by chess engines, it's called pruning (they'd be quite crippled without it). Maybe the neural network component is now in charge of it, but it's not a new thing.
- mvanaltvorst 6y agoOf course, I had accumulated all other pruning strategies under the "other optimisations".
- LittlePeter 6y agoLeela played stockfish 200 games and won with 106 - 94 [1]. Not sure which version of stockfish was used. Some of the Leela-Stockfish games are analyzed by agadmator on YouTube [2]. [1] https://www.chess.com/news/view/13th-computer-chess-championship-leela-chess-zero-stockfish https://www.chess.com/news/view/13th-computer-chess-champion... [2] https://www.youtube.com/watch?v=YtXZjKItuC8 https://www.youtube.com/watch?v=YtXZjKItuC8
- zone411 6y agoThis match was played before the NNUE version of Stockfish was introduced. Stockfish NNUE beat LC0 in TCEC season 19: https://www.chessprogramming.org/TCEC_Season_19 https://www.chessprogramming.org/TCEC_Season_19
- thomasahle 6y agoThe chess.com version is the old Stockfish from mid 2020. The NUE architecture was only put in place around August. If you see https://github.com/glinscott/fishtest/wiki/Regression-Tests https://github.com/glinscott/fishtest/wiki/Regression-Tests you'll notice that gave a very significant (100+ ELO) boost. There is a current tournament going on at tcec-chess.com/ which stockfish has been leading so far, but I see Leela has just caught up in the head to head. Of course Leela also keeps evolving.
- kohlerm 6y agoFYI Leela is in the lead ATM
- dmurray 6y agoIt's difficult to organise Leela-Stockfish as a fair fight, because they run on different hardware (CPU vs GPU) and both get substantial improvements by playing on better hardware. Traditionally this wasn't a big problem as every engine was more or less optimised for a fast Intel CPU with a moderate to large amount of RAM. The organisers would decree the specs of the championship hardware some time in advance. Now, (at least for TCEC, the other major engine tournament) they pick two hardware configurations, one CPU-heavy and one GPU-heavy, and give each team the choice. How do you balance those? It's been suggested you should make them equal in terms of watts of power, or dollar cost to buy, but neither of those are obviously best. In practice the TCEC organisers pick something close to what they picked last time but shade it against the winning engines, making the contest more even. Chess.com likely do something similar though they're less rigorous about the details.
- dan-robertson 6y agoIt seems that the strategy is to use the neural network to score various moves and then a search strategy to try to find moves that result in a favourable score. And this post focuses on some of the technical engineering details to design such a scoring network. In particular the scoring is split into two parts: 1. A matrix multiplication by a sparse input vector to get a dense representation of a position, and 2. Some nonlinear and further layers after this first step. And it seems that step 1 is considered more expensive. The way this is made cheap is by making it incremental: given some board s and output of the first layer b + Ws, it is cheap to compute b + Wt where a t is a board that is similar to s (the difference is W(t-s) but the vector t-s is 0 in almost every element.) This motivates some of the engineering choices like using integers instead of floats. If you used floats then this incremental update wouldn’t work. It seems to me that a lot of the smarts of stockfish will be in the search algorithm getting good results, but I don’t know if that just requires a bit of parallelism (surely some kind of work-stealing scheduler) and brute force or if it mostly relies on some more clever strategies. And maybe I’m wrong and the key is really in the scoring of positions.
- eutectic 6y agoI don't think the integer weights are neccessary for sparsity; they are just faster because they allow for low precision. Of course floats aren't strictly associative so you wouldn't get bitwise equivalence between the incremental and non-incremental updates, but I don't see how that would matter in this context.
- dan-robertson 6y agoThe point is that integer weights allow for incremental updates to be correct. If you used floats then they would drift away from correct as you applied more incremental updates. And neural networks can be quite sensitive to small errors
- eutectic 6y agoWith 32 bits of precision I don't see it being a big problem over a maximum of say 50 moves. Neural networks which are not overfit should be tolerant of a small amount of random (not crafted) noise.
- pcwelder 6y agoUsing a previous Stockfish scorer they trained an NN without any labelling effort. This is also similar to how unsupervised translation is done in some methods. They start from word->word dictionary results and iteratively train lang1->lang2 and lang2->lang1 models feeding on each other's output.
- brilee 6y agoIronically, a lot of the tricks Stockfish is using here are reminiscent of tricks that were used in the original AlphaGo and later discarded in AlphaGoZero. In particular, the AlphaGo paper mentioned four neural networks of significance: - a policy network trained on human pro games. - a RL-enhanced policy network improving on the original SL-trained policy network. - a value network trained on games generated by the RL-enhanced policy network - a cheap policy network trained on human pro games, used only for rapid rollout simulations. The cheap rollout policy network was discarded because DeepMind found that a "slow evaluations of the right positions" was better than "rapid evaluations of questionable positions". The independently trained value network was discarded because co-training a value and policy head on a shared trunk saved a significant amount of compute, and helped regularize both objectives against each other. The RL-enhanced policy network was discarded in favor of training the policy network to directly replicate MCTS search statistics. The depth and branching factor in chess and Go are different, so I won't say the solutions ought to be the same, but it's interesting nonetheless to see the original AlphaGo ideas be resurrected in this form. The incremental updates are also related to Zobrist Hashing, which the Stockfish authors are certainly aware of.
- mattalex 6y agoIt's also different because Stockfish uses Alpha-beta treesearch instead of MCTS: MCTS tries to deal with an explosion in search-space by only sampling very small parts of the searchspace and relying on a very good heuristic to guid that search process. In this case it is crucial to find the most relevant subset of the tree to explore it, so spending more time on your policy makes sense. Alpha-beta pruning however always explores the entire searchtree systematically (up to a certain depth using itd) and prunes the searchtree by discarding bad moves. In this case you don't need as good of an evaluation function because you search the entire tree anyways. Rather you need the function to be fast, as it is evaluated on many more states. In general AB-pruning only needs the heuristic for estimating the tail-end of the tree and for sorting the states based on usefulness, while MCTS uses all the above plus guiding the whole search process. Spending tons of time on the heuristic is not useful as even the optimal search order can only double your searchdepth. Don't get me wrong that still a lot (especially considering exponential blowup) but MCTS can surpass this depth easily. The disadvantage is that MCTS loses a lot of the guarantees of AB-pruning and tends to "play down to his opponent" when trained using self-play because the exploration order is entirely determined by the policy.
- glinscott 6y agoIf anyone wants to experiment with training these nets, it's a great way to get exposed to a nice mix of chess and machine learning. There are two trainers currently, the original one, which runs on CPU: https://github.com/nodchip/Stockfish https://github.com/nodchip/Stockfish, and a pytorch one which runs on GPU: https://github.com/glinscott/nnue-pytorch https://github.com/glinscott/nnue-pytorch. The SF Discord is where all of the discussion/development is happening: https://discord.gg/KGfhSJd https://discord.gg/KGfhSJd. Right now there is a lot of experimentation to try adjusting the network architecture. The current leading approach is a much larger net which takes in attack information per square (eg. is this piece attacked by more pieces than it's defended by?). That network is a little slower, but the additional information seems to be enough to be stronger than the current architecture. Btw, the original Shogi developers really did something amazing. The nodchip trainer is all custom code, and trains extremely strong nets. There are all sorts of subtle tricks embedded in there as well that led to stronger nets. Not to mention, getting the quantization (float32 -> int16/int8) working gracefully is a huge challenge.
- pk2200 6y agoJust wanted to say thanks for many years of fantastic work on both Stockfish and Leela. The computer chess community owes you a huge debt of gratitude!
- EvgeniyZh 6y agoInteresting how before A0 it was mainly "search matters the most", with crazy low branching factors to get deeper. It seems that humans were just better in search heuristics than in evaluation ones.
- billiam 6y agoAnalyses like this are a great indication that the main effect of refining neural networks around chess will make neural networks more exciting and ultimately make chess more boring.
- gelert 6y agoWhy do you believe that neural networks being better at chess would make it more boring? Genuinely I just don't follow the logic.
- spiantino 6y agoGreat writeup! One thing I don't understand is why it would be smarter to augment the inputs with some of the missing board information - particularly the availability of castling. Even though this network is a quick-and-dirty evaluation, seems like there's room for a few additional bits in the input and being able to castle very much changes the evaluation of some common positions.
- mattnewton 6y agoI am not an expert (and don’t have any idea what I am talking about), but wouldn’t this be captured the same way other future positional advantages are when evaluating the next “level” of possible board positions? Ie, the fact that a piece can later give check, or castle, or anything else if you first move it here is a function of that next move’s resulting position score. This network is just to signal to the final search algorithm how good the board looks in a given spot on tree of possible moves.
- spiantino 6y agoYes, the overall algorithm will do something much more precise and comprehensive, so it's not like this scoring function needs to be perfect. Still, it matters in a quick evaluation so I'm surprised it isn't in there. A few examples come to mind where, for instance, white has a bishop and knight attacking f7 and it's black to move. If black castles nothing interesting happens and white is likely overextended. But if black cannot castle white will win a rook. I can make up a similar situation with en passant, and this does come up, but a lot less frequently. So anyway, surprised you wouldn't toss a few bits in there, at least the 2 for ability to castle by each player.
- zetazzed 6y agoThe gap between white and black piece performance is massive in these top engines if I'm reading it right. LCZero won 0/ lost 4 as black and won 24 / lost zero as white (with lots of draws)? I had no idea the split was so big now. Do human tournaments look like this too these days? (From https://tcec-chess.com/ https://tcec-chess.com/)
- dsjoerg 6y agoOverwhelmingly the result is a draw. White can win sometimes, and wins for black are very rare. It's similar for top grandmasters, but not quite as stark.
- bonzini 6y agoNot really---similar for sure, but not as bad for black. Consider that computer chess tournaments do not start with the pieces in their initial positions. The first 5-15 moves are predetermined by human master players to achieve positions that are imbalanced while not having a side with a clearer advantage. Otherwise it would be even more of a draw fest. This makes things a bit worse for black, typically, because playing for the win as black means making some positional concession (that white can exploit after successfully defending) in exchange for the initiative and a chance for an attack. In human play, black has a little more ability than white to steer the game towards their opening preparation[1]. If they want to win, they will try to get have a position that they know in and out from (computer-assisted) home preparation. But there are still a lot of draws in classical (many hours per game) chess. [1] Chess openings that white chooses are typically called "attack" (King's Indian Attack) or "game" (Scotch Game), while those that black chooses are called "defense" (Sicilian defense). You'll find that there are a lot more "defenses" that black can choose from than "attacks".
- nl 6y agoThis is a really good article. Lots of pieces about neural network design skip over the design of the representation in the input stage, which is one of the key design issues when building a custom neural network. I love how much depth this articles goes into about that representation.