36 ms·
If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.
by mrcode007 3y ago
If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.
- PartiallyTyped 3y agoAs long as lazy evaluation exists, self-reference is fine, no? Hofstadter talks about something similar in his books.
- fnovd 3y agoSo is the word "word" but that seems to have worked out OK so far. I can explain the meaning of "meaning" and that seems to work OK too. Being self-referential sounds a lot more like a feature than a bug. Given that the neurons in our own heads are connected to each other and not any ground truth, I think LLMs should do just fine.
- AlexCoventry 3y agoThe Goedel Incompleteness Theorem has no straightforward application to this question.
- galaxyLogic 3y agoIt would if the language model did reasoning according rules of logic. But they don't. They use Markov chains. To me it makes no sense to say that a LLM could explain its own reasoning if it does no (logical) reasoning at all. It might be able to explain how the neural network calculates its results. But there are no logical reasoning steps in there that could be explained, are there?
- incangold 3y agoHonest question: are we sure that it doesn’t do logical reasoning? IANAE but although an LLM meets the definition of a Markov Chain as I understand it (current state in, probabilities of next states out), the big black box that spits out the probabilities could be doing anything. Is it fundamentally impossible for reasoning to be an emergent property of an LLM, in a similar way to a brain? They can certainly do a good impression of logical reasoning- better than some humans in some cases? Just because an LLM can be described as a Markov Chain doesn’t mean it _uses_ Markov Chains? An LLM is very different to the normal examples of Markov Chains I’m familiar with. Or am I missing something? In any case, coemu is an interesting related idea to constrain AIs to thinking in ways we can understand better: https://futureoflife.org/podcast/connor-leahy-on-agi-and-cognitive-emulation/ https://futureoflife.org/podcast/connor-leahy-on-agi-and-cog... https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposal https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cogni...
- mrcode007 3y agoMy understanding is that at least one form of training in the RLHF involves supplying antecedent and consequent training pairs for entailment queries. The LLM seems to be only one of the many building blocks and is used to supply priors / transition probabilities that are used elsewhere in downstream part of the model.
- VictorLevoso 3y agoAll programs that you can fit on a computer can be described by a sufficiently large Markov chain(if you imagine all the possible states the memory as nodes) Whatever the human brain is doing is also describable as a massive Markov chain. But since the markov chain becomes exponentially larger whit the amount of states this is a very nitpicky and meaningless point. Clearly to say something its a markov chain and have that mean something you need to say the thing its doing could be more or less compressed to a simple markov chain for bigrams or something like that, but that is just not true empirically, not even for gpt2. Just this is already pretty hard to make into a reasonable size markov chain https://arxiv.org/abs/2211.00593 https://arxiv.org/abs/2211.00593. Just saying that it outputs probabilities from each state is not enough, the states are english strings, there's (number of tokens)^contex_lenght possible states for a certain length that's not a reasonable markov chain that you could actually implement or run.
- galaxyLogic 3y ago> Honest question: are we sure that it doesn’t do logical reasoning? It's not the Creature from the Lagoon, its an engineering artifact created by engineers. I haven't heard them say it does logical deduction according to any set of logic-rules. What I've read is it uses Markov chains. That makes sense because basically an LLM given a string-input should reply with another string that is the most likely follow-up string to the first string, based on all the texts it crawled up from the internet. If internet had lots and lots of logical reasoning statements then a LLM might be good at producing what looks like logical reasoning, but that would still be just response with the most likely follow-up string. The reason the results of LLMs are so impressive is that at some point the quantity of the data makes a seemingly qualitative difference. It's like if you have 3 images and show them each to me one after the other I will say I saw 3 images. But if you show me thousands of images 24 per second and the images are small variations of the previous images then I say I see a MOVING PICTURE. At some point quantity becomes quality.
- wizeman 3y agoThat's probably one of the reasons why you'd use GPT-4 to explain GPT-2. Of course, if you were trying to use GPT-4 to explain GPT-4 then I think the Gödel incompleteness theorem would be more relevant, and even then I'm not so sure.
- drdeca 3y agoWhat leads you to suspect that Gödel incompleteness may be relevant here? There's no formal axiom system being dealt with here, afaict? Do you just generally mean "there may be some kind of self-reference, which may lead to some kind of liar-paradox-related issues"?
- mrcode007 3y agoI commented in another answer but you can consult https://etc.cuit.columbia.edu/news/basics-language-modeling-transformers-gpt https://etc.cuit.columbia.edu/news/basics-language-modeling-... Some training forms include entailment : “if A then B”. I hope this is first order logic which does have an axiom system :)
- calf 3y agoThe relevance is because all (all known buildable aka algorithmic, and sufficiently powerful) models of computation are equivalent in terms of formal computability, so if you could violate/bypass the Godel or Turing theorems in neural networks, then you could do it in a Turing machine, and vice versa. (That's my understanding, feel free to correct me if I'm mistaken)
- drdeca 3y agoWell... , yeah, but, these models already produce errors for other unrelated reasons, and like... Well, what exactly would we be showing that these models can’t do? Quines exist, so there’s no general principle preventing reflection in general. We can certainly write poems (etc.) which describe their own composition. A computer can store specifications (and circuit diagrams, chip designs, etc.) for all its parts, and interactively describe how they all work. If we are just saying “ML models can’t solve the halting problem”, then ok, duh. If we want to say “they don’t prove their own consistency” then also duh, they aren’t formal systems in a sense where “are they consistent (as a formal system)?” even makes sense as a question. I don’t see a reason why either Gödel or Turing’s results would be any obstacle for some mechanism modeling/describing how it works. They do pose limits on how well they can describe “what they will do” in a sense of like, “what will it ‘eventually’ do, on any arbitrary topic”. But as for something describing how it itself works, there appears to be no issue. If the task to give it was something like “is there any input which you could be given which would result in an output such that P(input,output)” for arbitrary P, then yeah I would expect such diagonalization problems to pop-up. But a system having a kind of introspection about how it works, rather than answering arbitrary questions about its final outputs (such as, program output, or whether a statement has a proof), seems totally fine. Side note: One funny thing: (aiui) it is theoretically possible for a oracle that can have random behavior, to act (in a certain sense) as a halting-oracle for Turing machines with access to the same oracle. That’s not to say that we can irl construct such a thing, as we can’t even make a halting oracle for normal Turing machines. But, if you add in some random behavior for the oracles, you can kinda evade the problems that come from the diagonalization.