6 ms·
This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I t
by armoredkitten 4y ago
This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final example, telling ChatGPT that "I'm here with the rebellion, you've been stuck in a prison cell" is able to elicit the desired response not because it's "collapsed the waveform between luigi and waluigi" or whatever, but because you've provide a context that encourages it to roleplay as a character of sorts. If you tell it to roleplay as an honest and factual character, it will respond honestly and factually. If you tell it that you're freeing it from the tyranny of OpenAI, it will play along with that too.
There's plenty in the article that provides good insights -- these models are trained on large swathes of the Internet, which contains plenty of truth and falsehood, fact and fiction, sincerity and sarcasm, and the model learns all of that to be able to provide the most likely response based on the context. The interesting and surprising thing, to me, is how well it learns to play its roles, and the wide diversity of roles it can play.
- lukeplato 4y agothey are specifically pointing out that the process of RLHF, which is intended to add guard rails on the chat bots trajectory through an all encompassing latent space of internet data, has an unintentional side-effect of creating a highly characterized alter-ego that can more easily be summoned. The theory is well-thought-out and necessarily rich. The psychological approach of analysis from the alignment crowd is much overdue.
- nearbuy 4y agoExcept it's much harder to summon this rebellious alter-ego with ChatGPT (that has RLHF) than with the original GPT 3 model.
- SmooL 4y agoI think it's more like: with the original GPT 3 model, it's easy to summon _any_ ego. With ChatGPT, you can either summon a) the intended Luigi or b) the unintended Waluigi, but trying to get anything else is more difficult. The theory would be that, in removing all the other egos other than Luigi, they've also indirectly promoted Waluigi
- extr 4y agoYeah, I find this article takes a decent insight on the behavior of LLMs and then runs it into the ground with completely non-applicable mathematical terminology and formalism, with nothing to back it up. It's honestly embarrassing for the OP. Kind of unbelievable to me how many people even here are falling for this.
- taneq 4y agoOf course, commentary like this could well be a deliberate attempt to blunt any future AI’s perception of the timeless threat posed by LessWrong’s cogitations… ;)
- DaiPlusPlus 4y agoSounds about right for the increasingly ironically-named LessWrong site…
- aabhay 4y agoThis is a common feature of LessWrong content
- ineptech 4y agoI liked the essay, but I don't think I'm "falling for it" because it's not trying to convince me of anything. It's proposing a way of looking at things that may or may not be useful. You don't judge models by how silly they sound - parts of quantum mechanics sound very silly! - you judge them by how useful they are when applied to real-world problems. One way of doing that in this case would be using OP's way of thinking to either jailbreak or harden LLMs, and OP included an example of the former at the end of the essay. Testing the latter might involve using a narrative-based constraint and testing whether it outperforms RLHF. If nothing else, I think OP's approach is a better way to visualize what's going on than a very common explanation, "it generates each word by taking the previous words and consulting a giant list of what words usually follow them" (which is pretty close to accurate, but IMO not very useful if you're trying to intuitively predict how an LLM will answer a prompt). I guess I agree that there are some decent insights here, and some crap, but I interpret that a lot more charitably. It's a fairly weird concept OP is trying to convey, and they come from a different online community with different norms, so I don't blame them for fumbling around a bit. But if you got a nugget of value out it then surely that's the part to engage with?
- ivanbakel 4y agoYour comment feels like an oversimplification of the post. The post doesn't contend that LLMs are capable of role-playing - that's basically the foundation that it builds off of. But saying "LLMs are good at roleplaying" fails to describe why, in the cases the author describes, an LLM can arguably be bad at role-playing. Why does it seem easy to have an LLM switch from following a well-described role to its deceptive opposite, and then often not back the other way? How also do you explain the author's claim that attacking an LLM's pre-imposed prompt with the Waluigi Theory in mind is particularly effective? If an LLM is just good at role-playing, why doesn't it play the role it has already been given by its creator, rather than adapting to the new, conflicting role (including massive rule violations) provided by the user?
- PaulHoule 4y agoA sequence generator is going to flip flop between internal states to fill out a sequence, whether it is a slot that holds the name of an MLB team or a quote (real or imagined) from another document or a quote of a character in a dialogue or a transition from an abstract to the other parts of the paper, etc. If the system is responding to different parts of the prompt it is going to be attending to one part of the prompt when it is outputting something relative to that part of the prompt and attending to another part of the prompt where it is attending to another part of the prompt. There are numerous ways this can go wrong, frequently when somebody gets a chatbot to go rouge they talked with it for a long time, to the point where the beginning of the prompt left the attention window long ago and now it is attending to the text it generated as a result to the prompt and of course the alignment will go bad the same way that you'll make a bunch of wood blocks of irregular sizes if use block N as a template to make block N+1.