2 ms·
The real issue with AlphaGo / MuZero is that it's still dramatically dumber than a dog, and not at all capable of understanding Go in the way a young child unde
by aithrowawaycomm 2y ago
The real issue with AlphaGo / MuZero is that it's still dramatically dumber than a dog, and not at all capable of understanding Go in the way a young child understands Go after playing a dozen games. These AIs are useful tools for playing Go, and for Go researchers to slice through the combinatorial mess, but they aren't intelligent in a very specific sense.
Consider the following thought experiment[1]: MuZero plays Donut Go, a variant of Go where a hole has been punched out of the middle of the board, but has not been retrained. Lee Sedol and a competent amateur would both be able to play competent amateur Donut Go by qualitatively adjusting their strategies to the new board in a "zero-shot" fashion: perhaps Lee would do much better, and the amateur makes a few boneheaded mistakes, but both would be competent.
MuZero would not be competent, it would make severe blunders over and over again because it's not capable of adjusting its strategy to the donut hole. It would need to be retrained. There's a categorical difference in flexibility, which is why the competent amateur is intelligent and MuZero is not.
[1] I am not sure if they've done a Donut Go experiment, but DeepMind did show that the same algorithm trained on Breakout could solve Level 1 with superhuman skill yet completely fail level 2. I am fairly confident my thought experiment holds water, but it would be interest to run it IRL. There is an issue in that I'm not sure how you would get MuZero to "see" the hole... if it has a deterministic check for move legality then perhaps that could be updated.
- ben_w 2y ago> MuZero plays Donut Go, a variant of Go where a hole has been punched out of the middle of the board, but has not been retrained. To be equivalent to a frozen/no retraining allowed AI, the human would need to have their brain chemistry prevented from altering synaptic weights. The option to freeze weights in AI models is the only way they can be safely productised. Or training runs fairly compared — didn't Kasparov allege that IBM cheated by changing the algorithm during the game? When it comes to "what even is this 'intelligence' thing?" there's definitely an argument about how much advantage these models take from the fact that silicon is around a million times faster than synaptic chemistry; but in practical terms, when it comes to "are we nearly obsolete?", it's the wall-clock time not the subjective time that matters — and AlphaZero took only 8 hours to beat its predecessor AlphaGo Zero, which itself beat AlphaGo Master, which in turn won 60:0 against professional players. So, how long would it take a human to get good at the new rules, and would it improve faster or slower than the wall-clock time of an AI learning the new constraints? > I am not sure if they've done a Donut Go experiment, but DeepMind did show that the same algorithm trained on Breakout could solve Level 1 with superhuman skill yet completely fail level 2. I am fairly confident my thought experiment holds water, but it would be interest to run it IRL. There is an issue in that I'm not sure how you would get MuZero to "see" the hole... if it has a deterministic check for move legality then perhaps that could be updated. Should be easy, given that MuZero did go, chess, shogi, and a standard suite of Atari games… But yeah, 2019 is an eternity in AI, should be easy to do what you suggest today.