3 ms·
Thanks for the thought-out reply. That said, we disagree on quite a few points. A 'lesson' is something to be learned from, and Rich Sutton explicitly mentions
by Matthyze 2y ago
Thanks for the thought-out reply. That said, we disagree on quite a few points. A 'lesson' is something to be learned from, and Rich Sutton explicitly mentions in his conclusion what we should learn from this lesson. But it's indeed not law. His argument is also not limited to ANNs, nor even ML. His first example concerns state space search and chess. In general, I think the field is moving further away from expert/domain knowledge and towards reliance on copious amounts of data and computation. End-to-end learning embodies this. LLMs incorporate essentially no domain knowledge about human language (processing). But of course it remains a matter of degree — certain domain knowledge remains very useful.
- YeGoblynQueenne 2y agoYes, I remember the argument. In computer chess the first AI system to dominate humans, IBM's DeepBlue, was based entirely on search and an opening book, so search + domain knowledge. Then AlphaGo which dominated in Go, was based on search, an opening book and a pair of self-playing deep neural nets and was equipped with the rules of Go. AlphaZero that followed, dropped the opening book but was still based on search and self-playing neural nets and was still given the rules of Go (and chess and Shoggi, if memory serves). And MuZero that followed that, dropped everything but the search and the self-playing neural nets. Now, Sutton argues that some of those systems at least did not rely on domain knowledge. They all did: Monte Carlo Tree Search, used for board game-playing AI agents, is nothing else but an encoding of domain knowledge - specifically, domain knowledge about the structure of two-player, complete information games. It's the same domain knowledge that was used to create the minimax-based DeepBlue software that won against Kasparov. Neural nets have still not managed to win against human players in board games without incorporating such a strong, knowledge-dependent component, as a game-tree search. So sutton is fudging the details. Not on purpose. He's an RL person. In RL, as in planning and other disciplines, folks tend to forget all the knowledge they put into their systems in the form of inductive biases, or auxiliary (but can't-do-without) algorithms like MCTS. He's like the proverbial fish that don't know what water is, because they swim in it.