11 ms·
LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we wi
by fnovd 3y ago
LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about ourselves?
- mrcode007 3y agoIf the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.
- PartiallyTyped 3y agoAs long as lazy evaluation exists, self-reference is fine, no? Hofstadter talks about something similar in his books.
- fnovd 3y agoSo is the word "word" but that seems to have worked out OK so far. I can explain the meaning of "meaning" and that seems to work OK too. Being self-referential sounds a lot more like a feature than a bug. Given that the neurons in our own heads are connected to each other and not any ground truth, I think LLMs should do just fine.
- AlexCoventry 3y agoThe Goedel Incompleteness Theorem has no straightforward application to this question.
- galaxyLogic 3y agoIt would if the language model did reasoning according rules of logic. But they don't. They use Markov chains. To me it makes no sense to say that a LLM could explain its own reasoning if it does no (logical) reasoning at all. It might be able to explain how the neural network calculates its results. But there are no logical reasoning steps in there that could be explained, are there?
- incangold 3y agoHonest question: are we sure that it doesn’t do logical reasoning? IANAE but although an LLM meets the definition of a Markov Chain as I understand it (current state in, probabilities of next states out), the big black box that spits out the probabilities could be doing anything. Is it fundamentally impossible for reasoning to be an emergent property of an LLM, in a similar way to a brain? They can certainly do a good impression of logical reasoning- better than some humans in some cases? Just because an LLM can be described as a Markov Chain doesn’t mean it _uses_ Markov Chains? An LLM is very different to the normal examples of Markov Chains I’m familiar with. Or am I missing something? In any case, coemu is an interesting related idea to constrain AIs to thinking in ways we can understand better: https://futureoflife.org/podcast/connor-leahy-on-agi-and-cognitive-emulation/ https://futureoflife.org/podcast/connor-leahy-on-agi-and-cog... https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposal https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cogni...
- mrcode007 3y agoMy understanding is that at least one form of training in the RLHF involves supplying antecedent and consequent training pairs for entailment queries. The LLM seems to be only one of the many building blocks and is used to supply priors / transition probabilities that are used elsewhere in downstream part of the model.
- VictorLevoso 3y agoAll programs that you can fit on a computer can be described by a sufficiently large Markov chain(if you imagine all the possible states the memory as nodes) Whatever the human brain is doing is also describable as a massive Markov chain. But since the markov chain becomes exponentially larger whit the amount of states this is a very nitpicky and meaningless point. Clearly to say something its a markov chain and have that mean something you need to say the thing its doing could be more or less compressed to a simple markov chain for bigrams or something like that, but that is just not true empirically, not even for gpt2. Just this is already pretty hard to make into a reasonable size markov chain https://arxiv.org/abs/2211.00593 https://arxiv.org/abs/2211.00593. Just saying that it outputs probabilities from each state is not enough, the states are english strings, there's (number of tokens)^contex_lenght possible states for a certain length that's not a reasonable markov chain that you could actually implement or run.
- wizeman 3y agoThat's probably one of the reasons why you'd use GPT-4 to explain GPT-2. Of course, if you were trying to use GPT-4 to explain GPT-4 then I think the Gödel incompleteness theorem would be more relevant, and even then I'm not so sure.
- drdeca 3y agoWhat leads you to suspect that Gödel incompleteness may be relevant here? There's no formal axiom system being dealt with here, afaict? Do you just generally mean "there may be some kind of self-reference, which may lead to some kind of liar-paradox-related issues"?
- mrcode007 3y agoI commented in another answer but you can consult https://etc.cuit.columbia.edu/news/basics-language-modeling-transformers-gpt https://etc.cuit.columbia.edu/news/basics-language-modeling-... Some training forms include entailment : “if A then B”. I hope this is first order logic which does have an axiom system :)
- calf 3y agoThe relevance is because all (all known buildable aka algorithmic, and sufficiently powerful) models of computation are equivalent in terms of formal computability, so if you could violate/bypass the Godel or Turing theorems in neural networks, then you could do it in a Turing machine, and vice versa. (That's my understanding, feel free to correct me if I'm mistaken)
- drdeca 3y agoWell... , yeah, but, these models already produce errors for other unrelated reasons, and like... Well, what exactly would we be showing that these models can’t do? Quines exist, so there’s no general principle preventing reflection in general. We can certainly write poems (etc.) which describe their own composition. A computer can store specifications (and circuit diagrams, chip designs, etc.) for all its parts, and interactively describe how they all work. If we are just saying “ML models can’t solve the halting problem”, then ok, duh. If we want to say “they don’t prove their own consistency” then also duh, they aren’t formal systems in a sense where “are they consistent (as a formal system)?” even makes sense as a question. I don’t see a reason why either Gödel or Turing’s results would be any obstacle for some mechanism modeling/describing how it works. They do pose limits on how well they can describe “what they will do” in a sense of like, “what will it ‘eventually’ do, on any arbitrary topic”. But as for something describing how it itself works, there appears to be no issue. If the task to give it was something like “is there any input which you could be given which would result in an output such that P(input,output)” for arbitrary P, then yeah I would expect such diagonalization problems to pop-up. But a system having a kind of introspection about how it works, rather than answering arbitrary questions about its final outputs (such as, program output, or whether a statement has a proof), seems totally fine. Side note: One funny thing: (aiui) it is theoretically possible for a oracle that can have random behavior, to act (in a certain sense) as a halting-oracle for Turing machines with access to the same oracle. That’s not to say that we can irl construct such a thing, as we can’t even make a halting oracle for normal Turing machines. But, if you add in some random behavior for the oracles, you can kinda evade the problems that come from the diagonalization.
- chrisco255 3y agoAre there any examples of an LLM developing concepts that do not exist or cannot be inferred from its training set?
- sgt101 3y agoThe training sets are so poorly curated we will never know...
- fnovd 3y ago"Cannot be inferred from its training set" is a pretty difficult hurdle. Human beings can infer patterns that aren't there, and we typically call those hallucinations or even psychoses. On the other hand, some unconfirmed, novel patterns that humans infer actually represent groundbreaking discoveries, like for example much of the work of Ramanujan. In a real sense, all of the future discoveries of mathematics already exist in the "training set" of our present understanding, we just haven't thought it all the way through yet. If we discover something new, can we say that the concept didn't exist, or that it "couldn't be inferred" from previous work? I think the same would apply to LLMs and their understanding of the way we encode information using language. Given their radically different approach to understanding the same medium, they are well poised to both confirm many things we understand intuitively as well as expose the shortcomings of our human-centric model of understanding.
- sebzim4500 3y agoIt is by definition impossible for an LLM to develop a concept that 'cannot be inferred from its training set'. On the other hand, that is an incredibly high bar.
- PeterisP 3y agoTautologically, every concept that anything (LLM, or human, or alien) develops can be inferred from the input data(e.g. training set), because it was.
- chrisco255 3y agoNo, it wasn't, language itself didn't even exist at one point. It wasn't inferred from training data into existence because such examples existed before. Now we have a dictionary of tens of thousands of words, which describe high level ideas, abstractions, and concepts that someone, somewhere along the line had to invent. And I'm not talking about imitation nor am I interested in semantic games, I'm talking about raw inventiveness. Not a stochastic parrot looping through a large corpus of information and a table of weights on word pairings. Has AI ever managed to learn something humans didn't already know? It's got all the physics text books in its data set. Can it make novel inferences from that? How about in math?
- ly3xqhl8g9 3y agoFirst of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw. With his left hand (controlled by his right hemisphere) he selected a shovel, which matched the snow scene. With his right hand (controlled by his left hemisphere) he selected a chicken, which matched the chicken claw. Next, the experimenter asked the patient why he selected each item. One would expect the speaking left hemisphere to explain why it chose the chicken but not why it chose the shovel, since the left hemisphere did not have access to information about the snow scene. Instead, the patient’s speaking left hemisphere replied, “Oh, that’s simple. The chicken claw goes with the chicken and you need a shovel to clean out the chicken shed”" [1]. Also [2] has an interesting hypothesis on split-brains: not two agents, but two streams of perception. [1] 2014, "Divergent hemispheric reasoning strategies: reducing uncertainty versus resolving inconsistency", https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522 https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522 [2] 2017, "The Split-Brain phenomenon revisited: A single conscious agent with split perception", https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf
- BaculumMeumEst 3y agothat is absolutely fascinating and also makes me extremely uncomfortable
- incangold 3y agoSame. We are so, so profoundly not what it feels like we are, to most of us anyway. I am morbidly curious how people are going to creatively explain away the more challenging insights AI gives us in to what consciousness is.
- ly3xqhl8g9 3y agoIt's probably way worse than we can imagine. Reading/listening to someone like Robert Sapolsky [1] makes me laugh I could have ever hallucinated about such a muddy, not even wrong concept as "free will". Furthermore, between the brain and, say, the liver there is only a difference of speed/data integrity inasmuch as one cares to look for information processing as basal cognition: neurons firing in the brain, voltage-gated ion channels and gap junctions controlling bioelectrical gradients in the liver, and almost everywhere in the body. Why does only the brain has a "feels like" sensation? The liver may have one as well, but the brain being an autarchic dictator perhaps suppresses the feeling of the liver, it certainly abstracts away the thousands of highly specialized decisions the liver takes each second solving adequately the complex problem space of blood processing. Perhaps Thomas Nagel shouldn't have asked "What Is It Like to Be a Bat?" [2] but what is it like to be a liver. [1] "Robert Sapolsky: Justice and morality in the absence of free will", https://www.youtube.com/watch?v=nhvAAvwS-UA https://www.youtube.com/watch?v=nhvAAvwS-UA [2] https://en.wikipedia.org/wiki/What_Is_It_Like_to_Be_a_Bat%3F https://en.wikipedia.org/wiki/What_Is_It_Like_to_Be_a_Bat%3F
- jmfldn 3y ago"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an LLM program how "it" knows something it will produce some text.
- nerpderp82 3y agoWhat if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer? What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?
- YawningAngel 3y agoWhat if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.
- int_19h 3y agoIf you direct the model output to itself and don't view it otherwise, how is it not an "internal stream of thought"?
- TeMPOraL 3y agoLLMs seem to me to be the "internal streams of thought". I.e. it's not LLMs that are missing an internal process that humans have, but rather it's humans that have an entire process of conscious thinking built on top of something akin to LLM.
- 3y ago
- elwell 3y agoAnd if the LLM is the explainer, it can lie to us if 'needed'.
- Sharlin 3y agoI’m sure LLMs are quickly going to learn to hallucinate (or let’s use the proper word for what they’re doing: confabulate) plausible-sounding but nonsense explanations of their thought processes at least as well as humans.
- Aerbil313 3y agoI don't think we invent new words just for AI to explain its thought process to us better. AI may explain more elaborately in our language instead.