4 ms·
If I understand correctly, you propose using GPUs so that traditional chess engines can search even more nodes per second. Leela evaluates three orders of magn
by 781 7y ago
If I understand correctly, you propose using GPUs so that traditional chess engines can search even more nodes per second.
Leela evaluates three orders of magnitude less nodes per second than Stockfish (70K vs 70M), yet it still won. You need to search smarter, not deeper, as this result shows.
- dragontamer 7y ago> You need to search smarter, not deeper, as this result shows. This result pits a GPU with 10,000 GOPS worth of compute and 500GBps of memory bandwidth against a CPU with only 100 GOPS of compute and 50GBps memory bandwidth. All I'm saying is: how do we test a "fair" comparison? Well... lets (somehow) rewrite the CPU algorithm to run on the GPU. Its an unsolved problem, maybe its completely unsolvable. But its at least a research-question worth pondering. --------- The TL;DR is: I think its possible to port Stockfish's evaluation functions over to a GPU. However, there are numerous unsolved research problems that need to be solved if this were to happen. The hash-table needs to be reworked, the search tree needs to be reworked, "Lazy SMP" won't work on a GPU, etc. etc. Most of the "tree search" methods need to be completely rewritten from scratch to be GPU-based. But the fundamental evaluation functions of Stockfish are brutally simple and look like they could work on a GPU shader to me.
- 781 7y agoStockfish tests improvements by also playing random games: > Changes to game-playing code are accepted or rejected based on results of playing of tens of thousands of games on the framework against an older "reference" version of the program https://en.wikipedia.org/wiki/Stockfish_(chess)#Fishtest https://en.wikipedia.org/wiki/Stockfish_(chess)#Fishtest Think about it, a human makes a small tweak, then they run a lot of games with it and decide if it's worth keeping or not. That's basically a terrible implementation of a neural network's gradient descent. Now imagine instead of a human coming up with a small tweak, you randomly search for it and take a holistic view of the board instead of micro-heuristics. Oh, you just invented AlphaZero :) BTW, since chess is highly parallel, you can run Stockfish on a computer cluster. Using this approach you can equalize the GPU power and have a balanced computer power match (of course, it will be totally unbalanced power-consumption wise). I'm not convinced Stockfish will win with equal computing power.
- dragontamer 7y ago> Now imagine instead of a human coming up with a small tweak, you randomly search for it and take a holistic view of the board instead of micro-heuristics. Oh, you just invented AlphaZero :) I'm fairly certain that the human brain does not work off of FP16 floating point numbers with back-propagated errors being calculated with differential equations to set the weights of our individual neurons. :-) Artificial neural networks are fascinating self-learning machines. But remember: they're artificial. There's nothing "human" about LeelaZero, AlphaGo, or any other CNN. Especially because AlphaGo / LeelaZero are augmented with an exceptionally powerful MCTS search functionality (no human counts the number of positions they visit and "balances" each node... nobody does that. MCTS is an extremely powerful computer algorithm for search, also an artificial construct) > I'm not convinced Stockfish will win with equal computing power. Ehh? The results are 10-Leela / 8-Stockfish / 82 draws. It was an exceptionally close set of 100 games. Note that Stockfish works with a global hash-table of chess positions it shares between threads. This methodology works with a "low" number of threads (ie: 16 threads, maybe even 64 threads). But it absolutely will not work at GPU-scale (~16,386+ SIMD threads on Vega64, or similar GPUs). It is not going to be an easy job to "port" Stockfish properly to a GPU-based system, or even to a cluster of 100x racked up computers. How do you efficiently share a global hash table across 100x clustered computers? I mean, you simply cannot. Stockfish simply isn't designed to scale that high. Stockfish is innately a single-node design that's constrained by the RAM. --------- The fact of the matter is: modern systems need a higher-form of scaling. I think LeelaZero / AlphaZero are "cheating", in that they've found methodologies that allow the huge amount of GPU easily be used. I think this is a wakeup call: that algorithms need to start looking at the GPU more carefully. Heavy compute definitely needs to start thinking about how to scale to 16,000+ threads and work on GPU-like systems.
- 781 7y agoYou keep on missing the point. Today's Stockfish running on a low end computer would still beat Stockfish from many years ago running on a much more powerful one, because it's evaluation function is significantly better. Chess is not strictly about computing power, and neural-network evaluation functions are vastly better.
- mot524 7y agoImplementing Stockfish's evaluation function on a GPU is almost certainly possible, but stupid. Stockfish is able to evaluate several million nodes per second. By the time it identified a position to evaluate and copied it to the GPU, it could have already finished evaluating it on the CPU. But a GPU can evaluate many positions in parallel... great... not helpful here because you don't know which positions you need to evaluate. This is a case of apples and oranges. CPUs are good at running conventional chess engines and GPUs aren't, that's just a fact and there's no way around it. GPUs are good at running some stuff like neural networks, so, great. If you need to run a neural network engine (LC0) then run it on a GPU, if you need to run a conventional engine, then run it on a CPU. Which is what people are doing now. Trying to change this is nonsensical, and trying to argue about which has more compute power makes as much sense as arguing about whether a car can drive on water better than a boat can drive on land.