5 ms·
The argument would be that that conceptual model is encoded in the intermediate-layer parameters of the model, in a different but analogous way to how it's enco
by GeneralMayhem 2y ago
The argument would be that that conceptual model is encoded in the intermediate-layer parameters of the model, in a different but analogous way to how it's encoded in the graph and chemical structure of your neurons.
- abernard1 2y agoI agree that's an argument. I would contend that argument is obviously false. If it were true, LLMs could multiply scalar numbers together trivially. It should be the easiest thing in the world for them. The network required to do that well is extremely small, the parameter sizes of these models are gigantic, and the textual expression is highly regular: multiplication is the simplest concept imaginable. That they cannot do that basic task implies to me that they have almost no conceptual understanding unless the fit is almost memorizable or the space is highly regular. That LLMs can't multiply numbers properly isn't surprising if they don't really understand concepts prior to emitting text. Where they do logical tasks, that can be done with minimal or no understanding, because syllogisms and logical formalisms are highly structured in text arguments.
- GaggiX 2y agoMultiplication requires O(n^2) complexity with the usual algorithm used by humans, LLMs have a constant amount of computation available and they are not really efficient machines for math evaluation. They can definitely evaluate unseen expressions and you train a neural network to learn how to do sums and multiplications, I have trained models on sums and they are able to do sums never seen during training, the model learns the algorithm just by giving it inputs and outputs.
- jdietrich 2y agoLLMs do contain conceptual representations and LLMs are capable of abstract reasoning. This is trivially provable by asking them to reason about something that is a) purely abstract and b) not in the training data, e.g. "All floots are gronks. Some gronks are klorps. Are any floots klorps?" Any of the leading LLMs will correctly answer questions of this type much more often than chance.
- LetsGetTechnicl 2y agoThat is not an example of a LLM being capable of abstract reasoning. Changing the question from "What is the capital of United States?" which is easily answerable to something completely abstract and "not in the training model" doesn't change that LLM's are just very advanced text prediction, and always will be. The nature of their design means they are incapable of AGI.
- jdietrich 2y agoThe question I gave is a literal textbook example of abstract reasoning. LLMs are just very advanced text prediction, but they are also provably capable of abstract reasoning. If you think that those statements are contradictory, I would encourage you to read up on the Bayesian hypotheses in cognitive science - it is highly plausible that our brains are also just very advanced prediction models.
- nsagent 2y agoYou're quite right that LLMs can seemingly do some abstract reasoning problems, but I would not say they aren't in the training data. Sure, the exact form using the made up word gronk might not be in the training data, but the general form of that reasoning problem definitely exists, quite frequently in fact.
- jdietrich 2y agoYes, but the general form of the problem tells you nothing about the answer to any specific case. To perform any better than chance, the model has to actually reason through the problem.
- cgag 2y agoHave you seen this? ``` You will be given a name of an object (such as Car, Chair, Elephant) and a letter in the alphabet. Your goal is to first produce a 1-line description of how that object can be combined with the letter in an image (for example, for an elephant and the letter J, the trunk of the elephant can have a J shape, and for the letter A and a house, the house can have an A shape with the upper triangle of the A being the roof). Following the short description, please create SVG code to produce this (in the SVG use shapes like ellipses, triangles etc and polygons but try to defer from using quadratic curves). ``` ``` Round 5: A car and the letter E. Description: The car has an E shape on its front bumper, with the horizontal lines of the E being lights and the vertical line being the license plate. ``` Image generated here: https://imgur.com/a/Ia4Q2h3 https://imgur.com/a/Ia4Q2h3 How does it "just" predict the letter E could be used in such a way to draw a car? How does it just text predict working SVG code that draws the car made out of basic shapes and the letter E? I don't know how anyone could suggest there are no conceptual models embedded in there.
- deleted 2y ago[deleted]
- Zambyte 2y ago> If it were true, LLMs could multiply scalar numbers together trivially. FWIW most large models can do it better than I can in my head.
- famouswaffles 2y ago>If it were true, LLMs could multiply scalar numbers together trivially. I mean, it's not like GPT-4 can't do this with more accuracy than a human without a calculator.
- nsagent 2y agoUsing Occam's razor, that is less probable than the model picking up on statistical regularities in human language, especially since that's what they are trained to do.
- mitthrowaway2 2y agoThat's hard to conclude from Occam's razor here. Or, "statistical regularities" may have less explanatory power than you think, especially if the simplest statistical regularity is itself a fully predictive understanding of the concept of temperature.