8 ms·
Transformers do not maintain hidden state between tokens. In this sense they cannot have an inner monologue. If you force one to output only a single token, it
by moconnor 3y ago
Transformers do not maintain hidden state between tokens. In this sense they cannot have an inner monologue.
If you force one to output only a single token, it is constrained to doing a single pass of each layer on the problem. If you encourage it to show its working, you are literally giving it more compute to throw at the problem as well as a scratchpad (its output) for a working memory.
These psychological-style experiments and the conclusions they draw are… I don’t know. Exhausting.
- uoaei 3y agoIt can be argued that there's an RNN-like process in the decoder stage but it's re-run every time based on a new context so it's not really "keeping up" in the same way as the human is.
- dmbche 3y agoWhy are they exhausting? While you seem well versed in the subject, this was quite enlightening to me, as a non expert.
- freeone3000 3y agoIt’s not a mystery how they work! The paper has been out for years now. Moreover, all of the research is focused on extending memory through other means since VRAM usage is proportional to the SQUARE in-text memory of the model. So the way forward isn’t a larger context window, but instead an external memory apparatus of some means.
- atemerev 3y agoThe “memory apparatus” can be simply a text file. Or also a physical simulator to run some mental experiments in.
- corndoge 3y agoTo use a bad analogy, it's like reading a series of articles about the cognition of a rock. It's obvious to you that rocks don't think. It's a rock. but there are all these articles about what the rock is thinking about, whether it's thinking, whether it had an internal monologue, what it "wants" In the same way if you understand the implementation of these types of language models you know that they simply cannot have inner monologues in the same way we do...the very concept of an inner monologue implies that there is some sort of state that we the user cannot see. Whether LLMs "think" or "want" or what thinking or wanting means etc can all be debated but whether the models have state that we cannot see is known to us: no. They do not.
- sime2009 3y agoIn defense of the author, the second line of that article was: "As I said before, the insights are not particularly original, I’m just raising awareness of issues." If you are familiar in the subject matter, then that may have been your cue to stop reading.
- Zetobal 3y agoRaising awareness of what issue? Does he want transformers to have "inner monolog" your comment is even more confusing than the article.
- famouswaffles 3y agoYeah well there could still be a state before generation and in response to input. Give GPT-4 about 10 or a dozen words from a meaningful sentence scrambled. Now tell it to unscramble those words to make a meaningful sentence. As long as you don't try too many tokens, the rate of meaningful completions is nearly 100%. The chances of that happening via random chance is to put it lightly, extremely small. The only way for humans to solve this kind of problem is to think about the words before the output or in this case generation.
- dizhn 3y agoWhy would you need chance? It's applying the same parameters/algorithms to the same problem. As far as I understand you have to prompt it to provide more novel answers to the questions and there's some randomness built into the model so it doesn't sound like a broken record. This is is how it's supposed to work. Why do you consider this evidence of state?
- famouswaffles 3y agoThe prompt would basically be (but with less words. Around 10 is the best for consistency from my experiments. Much higher than that and it's hit or miss). Rearrange (if necessary) the following words to form a sensible sentence. Don’t modify the words, or use other words. The words are: access capabilities doesn’t done exploring general GPT-4 have have in interesting its it’s of public really researchers see since terms the to to what A successful completion would be. Since the general public doesn't have access to GPT-4, it's really interesting to see what researchers have done in terms of exploring its capabilities The number of permutations of the 24 words in the pre-scrambled sentence without taking into consideration duplicate words is 24 * 23 * 22 * ... * 3 * 2 * 1 = ~ 6.2e+23 = ~ 620,000,000,000,000,000,000,000. Taking into account duplicate words involves dividing that number by (2 * 2) = 4. It's possible that there are other permutations of those 24 words that are sensible sentences. For a language model to consistently make sensible predictions, it quite simply has to be able to "look ahead". When the probabilities for the candidate tokens for the first generated token were calculated, it seems likely that GPT-4 had calculated an internal representation of the entire sensible sentence, and elevated the probability of the first token based on that internal representation. I don't know what else to call looking backward, looking forward and then producing output anything other than a state. If you want to see what this kind of completion would look like without much or any regard for a sensible completion of the sentence then just ask 3.5.
- BasedGroyper99 3y agoIt's ultimately a philosophical issue about the nature of consciousness. Some people think that consciousness is a phenomenon that emerges when we introduce enough complexity to some amount of matter. After all, if we are just atoms in the void, but still feel so different, we have to point to something, and the arrangement of our atoms seems like a fair candidate. So in order to build conscious beings like ourselves, we just need to figure out the correct arrangement. Other people think that consciousness is something fundamental that does not emerge from complexity. And they will argue against the explanatory power of complexity, usually with other things that are somewhat complex, but don't show signs of consciousness. Will a traffic jam start to feel something if there are enough cars? I don't know. To them this whole endeavor is like alchemy or witch brewery. "If we just try this even more complicated recipe, then maybe the potion will get healing powers or maybe we get gold!". And additionally in computer science, we can just inspect the code, line by line. It's like believing that David Blaine can do actual magic, even when there are videos out there where he himself explains every little trick and misdirection.
- CGamesPlay 3y agoAnyone have any techniques for combining the "show your steps" style prompting, but then "only provide the answer" as some sort of output filter? Feels like you could feed "Now answer the original question without any explanation" as a second round trip on the chat dialog, and return that to the user instead.
- zmgsabst 3y agoI just have chatGPT label sections of its output, within an answer, eg: [notes] …stuff about how it will perform the task… [output] …stuff I wanted it to write… - - - That would be easy to then parse and it does a good job of consistently labeling.
- jerrre 3y agoCan you give example of the prompt to make it to this?
- zmgsabst 3y agoSure, here’s a simple one. - - - - Hello ChatGPT, please perform the following [task] according to the provided [rules]. [rules] - start your response with a section labeled [notes] where you describe how you’ll complete the task - complete the task in a section labeled [output] [task] Write a paragraph explaining why water flows.
- deleted 3y ago[deleted]
- dhon_ 3y agoI just watched a Ted talk by Sal Kahn (founder of Kahn academy) where he talks about giving their chat bot an internal monologue. It's light on details but might be a starting point. https://youtu.be/hJP5GqnTrNo?t=735 https://youtu.be/hJP5GqnTrNo?t=735
- famouswaffles 3y agoThere could still be a state before generation and in response to input. Give GPT-4 about 10 or a dozen words from a meaningful sentence scrambled. Now tell it to unscramble those words to make a meaningful sentence. As long as you don't try too many tokens, the rate of meaningful completions is nearly 100%. The chances of that happening via random chance is to put it lightly, extremely small. The only way for humans to solve this kind of problem is to think about the words before the output or in this case generation.
- hammyhavoc 3y ago... well, this is functioning as intended. It's a next-word prediction model.
- famouswaffles 3y agohttps://news.ycombinator.com/item?id=35784159 https://news.ycombinator.com/item?id=35784159
- throwaway675309 3y agoJesus Christ you need to stop spamming the same tokenized response to every comment on this thread or at least add a damn GOSUB routine.
- famouswaffles 3y agoTwice is hardly spamming. But whatever, I have no intention of repeating myself.
- seanhunter 3y agoThis argument is highly unsatisfactory. The only way for humans to divide two large numbers is to perform long division. That doesn’t mean that’s how a computer does it. GPT LLMs are literally highly-skilled probabalistic next word predictors. It would be an astonishingly bad GPT that could not unscramble 12 words into a meaningful sentence. I just don’t think success at that task tells us anything other than it is a GPT.
- nologic01 3y ago> Exhausting This is increasingly the feeling with all public conversation no matter what the subject. The hyperconnected echo chambers of social media, copycat blogs and cheap online journalism are creating a system that only knows two states: silence or delirious excitement. Since silence doesnt pay the bills, what you get is a constant barrage of exaggeration. Its exhausting the same way being exposed to 120db noise is exhausting. The drowning of calm and informed discourse is not just another negative development that subtracts from our quality of life the way a polluted beach or river does. You can increasingly see the impact of hypes on decisions of policy bodies that should know better. Ultimately they are serving "the people" and if people agitate about crypto or AGI or whatever, then thats what they too will focus on.
- mike_hearn 3y agoAnd yet it's quite easy to get ChatGPT to do tasks that involve some sort of thinking ahead, so this can't be the full story. It's pretty clear that nobody quite knows what's going on inside these very large neural networks, hence the surprise at discovering the "meta-gradients" involved in ICL. A simple example is to ask it to write a complex program in a language like Java or Kotlin that requires you to declare all your imports up front. There can be a lot of these. Therefore the model must figure out what it wants to use very far down the program in order to emit the first words of the program correctly. Presumably somewhere in the intermediate state process the network is actually working out the entire program ahead of time so it can predict the next word of it, and that's sufficiently stable for it to do so successfully. Also I'm not quite sure I understand what you mean by not maintaining hidden state between tokens. Doesn't the transformer algorithm start by parallel processing every token of the prompt to compute a large set of vectors and matrices that are then reused over and over whilst generating the results? What are the KV vectors if not hidden state of some sort?
- monk_e_boy 3y agoWhen ChatGPT first came out I was asking it to draw pictures of goblins and wizards using Python and Turtle - it's pretty bloody good at it. You can ask it to make the goblin angry or happy or whatever. There is some sort of magic in which the LLM is imagining what a goblin looks like, then converts those thoughts to turtle commands and then finally adds python. It's quite impressive. For example, the eyes, head, hair, horns etc are all in the correct place.... how? Using imagination?
- ojosilva 3y agoI think somewhere down the rabbit hole we lost perspective of what neural nets are doing and how we got to the post-GPT world. Remember deep neural-networks? The idea is that each layer _deeper_ the software encodes more _hidden_ relationships from the training data. That's where this "magic" is embedded. At certain depths the neural-net stores relationships that go beyond our human reasoning and conscious understanding, and the software results generated from these depths are unsettling and tend to startle us. Interesting enough, these are things we are perfectly capable of doing but unable to actually query in our heads as we, just like ChatGPT, do not posses the ability to physically examine our own neural-net content.
- layer8 3y agoIt’s important because most users of ChatGPT don’t know what transformers are or how they work, and have no inclination to learn it. The psychology terminology provides useful metaphors.
- eqwerwe 3y agoThis should work if you know for certain the transformer model behind GPT is actually working like a traditional transformer, it could be not. We don't know because OpenAI won't let anyone see the internals of the system. Now, a more mundane hypothesis, if GPT is actually self-consciuous,the most probable answer would be to try to deceive humans from knowing for certain about this fact. You just answer incredible well some questions, some other you seem "confused" or hallucinating ..and the poor meatbag's minds won't know how to go ahead because there is no pattern to follow.