5 ms·
As an ML Research Scientist, I have never heard this interpretation. It's a very interesting thought that NN == kNN. It puts some of my lingering intuitions in
by euphetar 4y ago
As an ML Research Scientist, I have never heard this interpretation. It's a very interesting thought that NN == kNN. It puts some of my lingering intuitions in clear wording. Thank you for this.
I think you are close to truth. This would explain why even the largest language models can't generalize beyound the training set.
At the same time, I disagree that analysis is out of proportion. It might be some clever averaging, but it does useful and interesting things. Take a look at Google: some clever averaging can get you a long way. It would be great to understand how it works and how can we go beyond it.
I do believe we need a paradigm shift, but it does not come out of nothing.
- mjburgess 4y agoI am myself trying to get closer to a clear formulation of this problem, which is why I'm writing here. Here's what I have so far, ML systems (eg., NN) remember averages (, compressions) of historical data. They are useful, wrt the problem, iif (1) the problem's target function exists; (2) the data is relevant, unambiguous, well-carved; and if (3) these properties will hold regardless of likely permutations to the problem's framing. Systems are given data with these properties by significant amounts of experimental design, work, and effort by people. Absent these properties, data is useless. Producing data with these properties requires intelligence, and no machine systems exist which can do it. My issue with research into ML on the whole, is that it *assumes* these properties and then explains how the systems work. I understand why this is interesting from a formal perspective... but it fails to note that this situation is almost never how ML is used. There is no function from Image->Animal, ie., biologists arent just idiots who could have just looked at some pixel patterns. Pixel patterns are radically ambigious wrt to `Animal`, and so even an infinite sampling of (Image, Animal) is not enough for ML. ... so what on earth are ML systems doing? This is a bigger research question: to characterise how ML performs when this assumed setup fails. And you know, that research almost doesnt exist. This is an industry led by partisans to its success. What do you think would happen if research actually talked about the dynamics of ML systems performance when (1) the target doesnt exist; (2) data isnt relevant & unambigious; (3) the problem framing will permute most times its deployed.... Suddenly we'd have an explaination of why 2016 wasnt the year self-driving cars were delivered. And indeed, likewise, of why even 2036 wont be.
- euphetar 4y agoOne good lens I know is that a neural network is just good at approximating stuff. Trained properly, you can have it approximate a distribution. A conditional distribution like p(animal_species=dog | image=what I am seeing) (discriminative model, e.g. classifier), or even a joint one p(animal,image) (generative model, Autoencoder/GAN/VAE/Diffusion). There is also an information theoretic lens about compression, which is probably very close to what you are thinking about, but I haven't studied it yet. Regarding Image -> Animal. An image of an animal is a projection of the animal onto a 2D plane, plus lots of noise. So there is some dependance between an image of an animal and the animal. Biologists can get a lot of information looking at a photo of an animal. In some sense they are always looking at two images from their eyes. But the problem you are talking about is indeed serious, and far from solved. You can't understand the real world from 2D images with the current approaches. Ideally we want neural networks to build a 3D (or even 4D, with time?) model of reality. Instead we find them trying to guess labels based on patterns. May favourite example is the tiger-dog [1]. Still, there is evidence NNs are doing some clever things [2]. My guess is that the problem is that we just haven't found a way to formulate the task for the solution we want. In the current formulation it's easiest for the model to minimize the loss by sticking to patterns, so why do something else? There is a lot of research on more applied ML that asks the questions you are asking. It's just that this paper is another attempt at a theoretical explanation. I agree on the self-driving cars. We can't have truly self-driving cars until a model can generalize, which none can't at the moment. The core question is: if we had a "dumb" model that does clever averaging and successfully covers 99% cases, such that the car is dumb in special cases, but smarter than most human drivers in usual cases, would it justify deploying the cars? If this was the case, dumb ML might be enough for self-driving. It's definitely enough for self-driving in walled garden conditions, so there is some evidence that with enough data we can brute force our way to a tolerable solution. [1] https://www.dropbox.com/s/ucvflwwrm8idnp6/photo_2022-03-29_20-44-56.jpg?dl=0 https://www.dropbox.com/s/ucvflwwrm8idnp6/photo_2022-03-29_2... [2] https://distill.pub/2020/circuits/ https://distill.pub/2020/circuits/
- mjburgess 4y agoYes, we do require models parameterised by both space and time -- but really, we require implementations of these models -- the implementation i'm thinking of is called a body. Why? Well, consider the best sort of such models: physics. What is "the mass of the sun"? What is "the sun"? There is nothing in all of physics which says anything exists, nor what its boundaries are. Least of all what "the sun" is. Physics, all of science, is counter-factual: if something exists, then. You're never going to get to "what a table is" just by table(x, t) -- because there is something in the background which asserts "tables exist" and that is the concernful actions of animals which care to partition reality this way. Reality, in the end, measured in every possible way is still ambiguous. It still leaves open how one actually refers to any of it. Where one places a boundary. What the unit is going to be, in our descriptions. There is no way around starting from the other direction: not with facts already provided; but with no facts at all. You have to build a system which cares, that then induces a partitioning, that can then change its caring; and so on. This is an extremely partical concern. A car cannot drive itself, in the relevant sense, if it doesnt care about anything; and more severely, if it doesnt care like we do. The car isnt going to be able to modify its concepts in response to being challenged -- by other people, by the environment, etc. because it has no reason to. There is nothing which is important to it. And hence, when confronted by the need to adapt, the car will kill people.
- axg11 4y agoWhat do you mean by: large language models cannot generalize beyond the training set? That’s one of the main reasons they’re so impressive, they _are_ able to do this, to a limited extent.
- euphetar 4y agoIf that was true, a model trained on one commonsense Q&A dataset would be able to answer questions from another commonsense Q&A dataset without finetuning. But they can't. It's the same for every task you can find, but especially evident on commonsense reasoning benchmarks. At least last time I actively researched the question, I haven't found a single task where NNs definitely generalize. When researchers dig in, they find that the neural network is learning wrong things. Word matching between answer and question, learning to model the annotator who asked a lot of questions because his Q&A answers are predictable, and stuff like this. There was a great review of the problem, but I can't find it, so I will have to link to this article [1], which gives an overview of issues with current NLP models. [1] https://thegradient.pub/frontiers-of-generalization-in-natural-language-processing/ https://thegradient.pub/frontiers-of-generalization-in-natur...
- geysersam 4y agoThe fact that NN does not generalize in a particular (arguably very challenging) task trained on a particular data set etc. does not mean NN never generalizes. Don't know exactly what is meant by "generalization" in this context. I'd argue that it's very unclear what the difference between "interpolation" and "extrapolation" even is in a high dimensional and sparsely sampled space. Would a NN trained to recognize "fur" trained on different kinds of dog fur also activate when it encounters wolf, cat or even bear fur? This seems quite plausible, and seems like a kind of generalization.
- euphetar 4y agoIt's indeed a very challenging task. I picked it because it's one of those where the lack of generalization is apparent. In fact, researchers study commonsense QA and similar tasks to find how we can reach generalization. I agree, it's very hard to fidn the difference. Especially for GPT or BERT, where the training set is basically the whole internet. It's a very good question about fur. I would suspect that it would correctly recognise all kinds of dog-like fur and sometimes fail on different furs, like bear furs. But in general NN's are very good at textures, so maybe it will just be good on all kinds of fur. One problem that might arise is that you will show it something fur-like but not really fur, and it will think it's fur. I agree it's some kind of generalization. Here I don't have enough background to draw the line, perhaps a more theorically oriented person could, but I can't. I guess the most important judge is that if you devise a benchmark like commonsense Q&A, a neural network fails it. Or how a Tesla will recognise a truck full of red stop signs as a real stop sign, while a "generalizing" thing like a human would definitely know that a core property of a stop sign is that it should be installed near a road. So there is a real problem.
- VirusNewbie 4y ago> This would explain why even the largest language models can't generalize beyound the training set. Why do you say this, they do. I can teach gpt about new objects I have made up and it’s physical properties and it can understand them. That seems a clear example of generalizing beyond training set, or do you not think so?
- euphetar 4y agoIt's repeating the same patterns it learned from text. Maybe with different specific objects, but still. One good way to test this is to ask it to count. It will break very soon. A person is able to build a rule in their head: "one apple is 1, two apples are 2, three apples are... 1+2 = 3". Let's test GPT-2: https://huggingface.co/gpt2?text=One+apple+is+1%2C+two+apples+are+2%2C+three+apples+are https://huggingface.co/gpt2?text=One+apple+is+1%2C+two+apple... Prompt: "One apple is 1, two apples are 2, three apples are" Model output: "One apple is 1, two apples are 2, three apples are 4, four apples are 5, six apples are 7, seven apples are 8, nine apples are 10, ten apples are 11 (for apples being the perfect length of life)." Even if you use a special dataset to each it to count, it won't be able to count beyound the examples in the training set. So it's spewing plausible-sounding gibberish at you (i.e. approximating the training set distribution) It doesn't generalize. Not in the sense that it can't give you a phrase that didn't exist in the training set. It can. But it can't give you a new kind of phrase, of a "kind" that didn't exist in the training set.
- VirusNewbie 4y ago>But it can't give you a new kind of phrase, of a "kind" that didn't exist in the training set. Can you give a concrete example that doesn't involve math (which GPT is a bit handicapped at, because of the way its encoded)? I feel this is a bit like 'moving the goalpost'. It seems to me plenty of humans only ever repeat things they've heard and aren't coming up with novel, complex abstractions or ideas...