19 ms·
Me and a friend were discussing PINNs, and he made an argument against them: The Bitter Lesson. PINNs are a way of incorporating domain knowledge into ML models
by Matthyze 2y ago
Me and a friend were discussing PINNs, and he made an argument against them: The Bitter Lesson. PINNs are a way of incorporating domain knowledge into ML models. The Bitter Lesson, for those unaware, is a famous essay by Rich Sutton that states that the history of AI is full of attempts of methods guided by domain/expert knowledge, but that ultimately, all such methods were overtaken by methods that simply scaled data and/or computation.
http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
I would love to hear HN's take on this argument.
- add-sub-mul-div 2y agoComputer science is currently subservient to an economic climate in which the only viable business is one that scales revenue without scaling labor. That's the bitter lesson.
- PaulHoule 2y agoBut isn't that the story of technology and civilization? Hunter-Gatherers produced no surplus and couldn't support a complex and unequal society. Early agriculture could support a pyramid, but not very high. It used to be almost everyone worked in agriculture, now about 1% does, so the others are free to do something else. Prior to the microprocessor making a computer required manual assembly of thousands of parts, early microprocessors contained thousands of parts manufactured by a small number of photographic and chemical steps, and the number of parts has grown into the billions without the number of steps expanding millions of times.
- rangestransform 2y agoImproving labour productivity is good, actually
- esafak 2y agoIt works when data is cheap, and model size is not an issue.
- nicoco 2y agoScaling data is not always possible. It's really hard to get your hands on good labelled medical imaging data, for instance. Maybe it makes sense to try to incorporate insights from biology and physiological instead of hoping that the neural net will "get it" from seeing enough data.
- constantcrying 2y ago>Maybe it makes sense to try to incorporate insights from biology and physiological instead of hoping that the neural net will "get it" from seeing enough data. For medical imaging your solver is most likely "physics based" in any case. PINNs want to encode the physics inside a neural network, instead of developing an appropriate algorithm (which obviously has to also incorporate the physics) which solves the problem and which can be analyzed mathematically.
- Archit3ch 2y agoThe Bitter Lesson only applies when you don't know the function you are modeling (e.g. perfect chess strategy or text-to-speech). On the contrary, neural networks will never give you a better way to convert to the frequency domain than the Fourier Transform. At best, they might approximate it.
- Matthyze 2y agoRight. And the use-case for PINNs is, e.g., modeling a known function that is not solvable analytically?
- constantcrying 2y ago>And the use-case for PINNs is, e.g., modeling a known function that is not solvable analytically? No. That is the use case for numerical analysis, where you can develop a high performance, accurate algorithm based on the mathematical analysis of the problem. It is really kind of silly to want to encode the solution in some neural network.
- Matthyze 2y agoSo, what's the use-case for PINNs, then?
- constantcrying 2y agoAllegedly you can use it to solve PDEs, although that does not seem to work well. Maybe there are some very special problems, without well developed silver where they work.
- YeGoblynQueenne 2y agoIt's not really an argument, it's more of an observation; and certainly not meant to be taken as a law, as in "thou shalt not use background knowledge even if you can". See, the reason for the Bitter Lesson being a lesson is that Neural Nets, that Sutton is mainly writing about, are pretty crap at representing background knowledge. The only way you can store expert knowledge in a neural net is to modify its structure and its weights. The weights you can only modify by some kind of learning procedure like backprop, in practice. Very limited forms of background knowledge, like convolutions can be encoded in a neural net's structure, but imagine trying to represent, I don't know, the last ten lines of code you wrote today as a bunch of neural net connections. Continuous functions is just not the right kind of notation for that sort of thing. If neural nets were any better at encoding background knowledge, they would use it, but they can't so they have to rely on data. And that's why they need so much of it. Background knowledge functions as a strong inductive bias- it directs the search for a hypothesis to hypotheses that we know make sense (again, think of convolutions). Without background knowledge, or with only a little background knowledge, you need tons of examples to learn anything useful. So the Bitter Lesson is basically making a virtue out of necessity. In any case, it's not a prescriptive thing, only descriptive.
- Matthyze 2y agoThanks for the thought-out reply. That said, we disagree on quite a few points. A 'lesson' is something to be learned from, and Rich Sutton explicitly mentions in his conclusion what we should learn from this lesson. But it's indeed not law. His argument is also not limited to ANNs, nor even ML. His first example concerns state space search and chess. In general, I think the field is moving further away from expert/domain knowledge and towards reliance on copious amounts of data and computation. End-to-end learning embodies this. LLMs incorporate essentially no domain knowledge about human language (processing). But of course it remains a matter of degree — certain domain knowledge remains very useful.
- YeGoblynQueenne 2y agoYes, I remember the argument. In computer chess the first AI system to dominate humans, IBM's DeepBlue, was based entirely on search and an opening book, so search + domain knowledge. Then AlphaGo which dominated in Go, was based on search, an opening book and a pair of self-playing deep neural nets and was equipped with the rules of Go. AlphaZero that followed, dropped the opening book but was still based on search and self-playing neural nets and was still given the rules of Go (and chess and Shoggi, if memory serves). And MuZero that followed that, dropped everything but the search and the self-playing neural nets. Now, Sutton argues that some of those systems at least did not rely on domain knowledge. They all did: Monte Carlo Tree Search, used for board game-playing AI agents, is nothing else but an encoding of domain knowledge - specifically, domain knowledge about the structure of two-player, complete information games. It's the same domain knowledge that was used to create the minimax-based DeepBlue software that won against Kasparov. Neural nets have still not managed to win against human players in board games without incorporating such a strong, knowledge-dependent component, as a game-tree search. So sutton is fudging the details. Not on purpose. He's an RL person. In RL, as in planning and other disciplines, folks tend to forget all the knowledge they put into their systems in the form of inductive biases, or auxiliary (but can't-do-without) algorithms like MCTS. He's like the proverbial fish that don't know what water is, because they swim in it.