10 ms·
The structure of having X apples in Y buckets is the same as the structure in the expression "X * Y", as long as the expression exists in a context that can par
by yldedly 5y ago
The structure of having X apples in Y buckets is the same as the structure in the expression "X * Y", as long as the expression exists in a context that can parse it using the rules of arithmetic, such as a human, or a calculator.
These language models lack context, not just for arithmetic, but for everything. They can't parse "X * Y" for any X and Y, they've just associated the expression with the right answer for so many values of X and Y, that we get fooled into thinking they know the rules.
We get fooled into thinking they've learned the structure of the world. But they've only learned the structure of text.
- whimsicalism 5y agoIt would be trivial for a network of this size to code general rules for multiplication. At a certain point, when you have enough data, finding the actual rule is actually the easier solution than memorizing each data point. This is the key insight of deep learning.
- yldedly 5y agoReally? Better inform all the researchers working on this that they're wasting their time then: https://arxiv.org/abs/2001.05016 https://arxiv.org/abs/2001.05016 More fundamentally, any finite neural net is either constant or linear outside the training sample,depending on the activation function. Unless you design special neurons like in the paper above, which solves this specific problem for arithmetic, but not the general problem of extrapolation.
- FeepingCreature 5y agoIsn't that per-layer?
- yldedly 5y agoNo, no matter how many piecewise linear functions you compose, the result is still piecewise linear.
- FeepingCreature 5y agoWell sure, but neurons are still universal approximators. Any CPU is a sum of piecewise linear functions. I don't see where this meaningfully limits the capabilities of an AI, since once we're multilayer there's no 1:1 relation between training samples and piece placement in the output.
- yldedly 5y agohttps://medium.com/analytics-vidhya/you-dont-understand-neural-networks-until-you-understand-the-universal-approximation-theorem-85b3e7677126 https://medium.com/analytics-vidhya/you-dont-understand-neur...
- FeepingCreature 5y agoI just don't see how that's relevant. Nobody uses one-hidden-layer networks anymore. Whatever GPT is doing, it has nothing to do with approximating a collection of samples by assembling piecewise functions, except in the way that Microsoft Word is based on the Transistor.
- yldedly 5y agoSounds like no amount of math will convince you otherwise.
- FeepingCreature 5y agoShould math about a vaguely related topic convince me about this? Multilevel ANNs act differently than one-level ANNs. Transformers simply don't have anything to do with the model of approximating functions by assembling piecewise functions. This is akin to arguing that computers can't copy files because the disjunctive normal form sometimes needs exponential terms on bit inputs, so obviously it cannot scale to large data sets - yes, that is true about the DNF, but copying files on a computer simply does not use boolean operations in a way that would run into that limitation. The way that Transformers learn has more to do with their multilayering than with the transformation across any one layer. Universal approximation only describes the things the network learns across any pair of layers, but the input and output features that it learns about in the middle are only tangentially related to the training samples. You cannot predict the capabilities of a deep neural network by considering the limitations of a one-layer learner.
- mjburgess 5y ago> any finite neural net is either constant or linear outside the training sample Hence why the structure of our bodies has to include the capacity for imagination. Our brain structure does not record everything that has happened. It permits is to imagine an infinite number of things which might happen. We do not come to understand the world by having a brain-structure isomorphic to world structure -- this is none-sense for, at least, the above reason. But also, there really isnt anything like "world structure" to be isomorphic to. Ie., brains arent HDDs. They are, at least, simulators. I dont think we'll find anything in the brain like "leaves are green" because that is just a generated public representation of a latent-simulating-thought. There isnt much to be learned about the world from these, they only make sense to us. That all the text of human history has associations between words is the statistical coincidence that modern NLP uses for its smoke-and-mirrors. As a theory of language it's madness.
- hackinthebochs 5y ago>We get fooled into thinking they've learned the structure of the world. But they've only learned the structure of text. To what degree does the structure of text correspond to structure of the world, in the limit of a maximally descriptive text corpus? Nearly complete if not totally complete, as far as I can tell. What is left out? The subjective experience of being embodied in the world. But this subjective experience is orthogonal to the structure of the world. And so this limitation does not prevent an understanding of the structure.
- yldedly 5y agoThe point is that not only is it impossible to infer the structure of the world from text, deep learning is incapable of learning about or even representing the world. The reason language makes sense to us is that it triggers the right representations. It does not make sense intrinsically, it's just a sequence of symbols. Learning about the world requires at least causal inference, modular and compact representations such as programming languages, and much smarter learning algorithms than random search or gradient descent.
- FeepingCreature 5y agoIt sounds like you're arguing that GPT doesn't work because it cannot work. However, it does work. So how does PaLM understand causal chains and explain jokes that it has never seen before?
- yldedly 5y agoIt doesn't. It's pattern matching, and you're seeing cherry picked examples. The pattern matching is enough to give the illusion of understanding. There's plenty of articles where more thorough testing reveals the difference. Here are two: https://medium.com/@melaniemitchell.me/can-gpt-3-make-analogies-16436605c446 https://medium.com/@melaniemitchell.me/can-gpt-3-make-analog... But you could also just try one of these models, and see for yourself. It's not exactly subtle. https://www.technologyreview.com/2020/08/22/1007539/gpt3-openai-language-generator-artificial-intelligence-ai-opinion/ https://www.technologyreview.com/2020/08/22/1007539/gpt3-ope...