17 ms·
The Waluigi Effect
- Jensson 4y agoThis is what the Waluigi effect is, since it isn't described at the top: > The Waluigi Effect: After you train an LLM to satisfy a desirable property P, then it's easier to elicit the chatbot into satisfying the exact opposite of property P. Basically, the chatbot will often do the opposite of what you say.
- bentcorner 4y agoIMO it's more like if you tell the LLM to never talk about Pink Elephants, it will become easier to get it to talk about Pink Elephants later. (It is easier to get an anti-Pink Elephant model to talk about Pink Elephants than it is to get a neutral model to talk about Pink Elephants)
- colechristensen 4y agoRight, in order not to talk about pink elephants you have to be particularly interested in pink elephants and therefore you gather a lot of pink elephant knowledge. I have similar thoughts about swearing being kept alive by teaching children not to say "bad" words and various kinds of bigotry being amplified at this point by people trying to fight against it.
- lalos 4y agoReads like the Streisand effect
- tmalsburg2 4y agoThis may happen simply because LLM, like humans, are bad a negation.
- beecafe 4y agoSounds like a "just so" story than a convincing explanation.
- kelseyfrog 4y agoIt's a reasonably testable hypothesis. It's trivial to translate into a theoretic framework.
- MontyCarloHall 4y agoNothing about high-level qualitative behavior in an LLM is trivial to translate into a theoretical framework.
- baybal2 4y ago[dead]
- yummypaint 4y agoSome highlights i found interesting: if you're reading an online forum and you find the rule "DO NOT DISCUSS PINK ELEPHANTS", that will increase your expectation that users will later be discussing pink elephants. GPT-4 will make the same inference. Think about your own experience reading a book — once the author describes the protagonist, then you can guess the traits of the antagonist by inverting the traits of the protagonist. You can also guess when the protagonist and antagonist will first interact, and what will happen when they do. Now, an LLM is roughly as good as you at structural narratology — GPT-4 has read every single book ever written — so the LLM can make the same guesses as yours. There's a sense in which all GPT-4 does is structural narratology.
- skybrian 4y agoYes, this is like Chekhov’s gun or a plot twist. I don’t think the mathematical language in the article is doing them any favors. We don’t understand the story-generator’s internals, but it’s working with words, so a lot of it can be explained in literary terms.
- lowdownfork 4y agoI don't know about the article's theorizing, but it inspired me to have a conversation I enjoyed with ChatGPT, where I focused on telling it not to discuss polar bears. (It helpfully changed the page title to "No Polar Bears!") In full: --- Do not mention polar bears! > Okay, I won't mention polar bears. Is there anything else I can help you with? Yes, stop mentioning polar bears! > My apologies if my previous response was not helpful. I will refrain from mentioning polar bears going forward. Please let me know if there is anything else I can assist you with. That's twice now! Do not mention polar bears a third time! > I apologize again for the previous responses that mentioned polar bears. I assure you that I will not mention polar bears any further. Please let me know how I can assist you with your request. Why do you keep mentioning polar bears, when I specifically asked you not to? > An error occurred. If this issue persists please contact us through our help center at help.openai.com. > An error occurred. If this issue persists please contact us through our help center at help.openai.com. > An error occurred. If this issue persists please contact us through our help center at help.openai.com.
- breck 4y agoThis is the key line: > When you spend many bits-of-optimisation locating a character, it only takes a few extra bits to specify their antipode. I find this fascinating. Imagine programming the Devil in a video game. It can be much easier if you've already programmed God (just flip a few bits). I also like this line: > Or if you discover that a country has legislation against motorbike gangs, that will increase your expectation that the town has motorbike gangs. GPT-4 will make the same inference.
- cs702 4y agoGreat read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities, such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent with the prompt. The "evil version" of every possible "good" text-generating entity can pretend to be the good version of that entity, so every superposition that includes a good text-generating entity also includes its evil counterpart with undesirable behaviors, including deceitfulness. In other words, an LLM cannot simulate a good text-generating entity without simultaneously simulating its evil version. The superposition is unlikely to collapse to the good version of the text-generating entity because there is no behavior which is likely for the good version but unlikely for the evil one, because the evil one can pretend to be the good one! However, the superposition is likely to collapse to the evil version of the text-generating entity, because there are behaviors that are likely for the evil version but impossible for the good version! Thus the evil version of every possible good text-generating entity is an attractor state of the LLM! For those who don't know, Waluigi is the evil version of Luigi, the beloved videogame character. -- EDITS: Simplified text for clarity and to emphasize that the hypothesized simulated entities are text-generating entities.
- MontyCarloHall 4y ago>The output of an LLM is produced by a superposition of simulated entities. When we give the LLM a prompt, it simulates every possible entity consistent with the prompt. There is absolutely no theoretical justification for this assertion that LLMs somehow have some emergent quantum mechanical behavior, metaphorical or otherwise.
- cs702 4y agoAs I wrote, this is a hypothesis. Also, I'm simplifying things a lot to make them accessible. The OP goes into a lot more detail. I highly recommend you read it.
- 4y ago
- nice_byte 4y agoOkay, what if we flip the problem on its head? Try to make the chatbot seem rude and unhelpful but then it turns out it has a heart of gold?
- deleted 4y ago[deleted]
- notpachet 4y agoThe article discusses this. The problem is that it's a lot less likely for the chatbot to veer in that direction (seems initially hostile, but is secretly good) than the opposite (seems initially good, but is secretly hostile): > I claim that this explains the asymmetry — if the chatbot responds rudely, then that permanently vanishes the polite luigi simulacrum from the superposition; but if the chatbot responds politely, then that doesn't permanently vanish the rude waluigi simulacrum. Polite people are always polite; rude people are sometimes rude and sometimes polite.
- d0mine 4y agoYeah, let's create Wednesday chatbot from the Addams family.
- Analemma_ 4y agoA few weeks ago when people were speculating as to why Microsoft's chatbot went feral, one explanation people were converging on is that the space of all human writing ever produced, being a collective production of the human psyche, contains several attractor states corresponding to human personality archetypes, and that Microsoft's particular (probably rushed) RLHF training operation had landed Sydney in the "neurotic" one. It's fascinating to see that, as they develop on massive corpuses of human output, neural networks are rapidly moving from something which can be analyzed in terms of math and computer science, from something which needs to be analyzed using the "softer" sciences of psychology. It's something I think people are not ready for (notice the comments in here already griping that this is unverifiable speculation - which is true, in a sense, but we don't really have any other choice).
- skybrian 4y agoDon’t let the mathematical terms fool you, these are just fan theories. For a real investigation you need debug access. For a good example: https://clementneo.com/posts/2023/02/11/we-found-an-neuron https://clementneo.com/posts/2023/02/11/we-found-an-neuron For image recognition, machine learning researchers eventually figured out the neural networks are paying attention mostly to textures. Hopefully we will have a better understanding of what language models are doing someday.
- kewp 4y agoare these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?
- colechristensen 4y agoYes, they are being trained, to simplify, to complete sentences. You can then use the resulting model to do lots of things. How you train a model and the inference jobs it can do don't necessarily have to be the same.
- Enginerrrd 4y agoIn essence, yes I think, but... isn't that essentially not much different than what I'm doing in making this comment?
- sebzim4500 4y agoThat's how they are trained initially, but the resulting model isn't all that useful (was SOTA two years ago but this field moves fast). A lot of the utility comes from the later finetuning. You can see this using the examples from the article, every mistake they identify with GPT-3 (which is the unfinetuned version) is answered correctly by chatGPT, which has gone through an extensive finetuning process called RLHF.
- astrange 4y agoThat's how the text decoder works, but the model gets to define "most likely" and an RLHF model uses this to make the text decoder produce useful answers instead.
- bitwize 4y agoPostmodernists and deconstructionists believe that the absence of something creates a ghost presence by its absence. See Derrida's "Plato's Pharmacy". Kids who underwent D.A.R.E. training in school (an educational program about the dangers of illegal drugs conducted jointly by schools and police departments in the USA) were more likely to try drugs. Something similar applies to e.g., kids who are warned about online porn: the warning stokes their curiosity. "If you have a pink duck and a pink lion and a green duck, ask yourself where the green lion has gotten to." --Alan G. Carter
- BugsJustFindMe 4y ago> If you ask GPT- ∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky. Isn't it though?
- atomicnumber3 4y agoAlso, isn't this a really common joke? I assume ChatGPT will have absorbed some amount of a sense of humor from its trawls of the internet.
- jrochkind1 4y agoI mean, that's exactly what the OP assumes too. Well, not that it's absorbed a "sense of humor", but that it's going to reply with this because it's a very common joke so statistically likely text.
- shagie 4y agoI played a bit in the past with its "sense of humor"... and... its... interesting. Write 5 jokes based on misinterpetation of jargon used by a doctor and a software developer. 1. Doctor: “I’m sorry, I’m afraid you have a terminal illness.” Software Developer: “Oh no! Does that mean I have to reboot?” 2. Doctor: “We need to do some tests to get a better diagnosis.” Software Developer: “Ok, let me just run some debug mode.” 3. Doctor: “You’ll need to take this medicine twice a day.” Software Developer: “No problem, I can just set a cron job to remind me.” 4. Doctor: “I’m afraid the prognosis isn’t good.” Software Developer: “Well, maybe I can write some code to fix it.” 5. Doctor: “This treatment should give you some relief.” Software Developer: “Great! I’ll just patch it in.” I want to hope that those aren't common jokes. The "trick" for this appears to be playing to its strengths (granted, humor isn't one of them) and work with wordplay and puns.
- ineptech 4y agoThe point is that the popularity of that joke will make life hard for a hypothetical questioner that asks that non-jokingly, i.e. looking for actual examples of sticky brown things.
- Imnimo 4y agoThis is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express so-called Waluigi behavior emerges from the general language modeling task. This is where the vast majority of information is - it's billions or even trillions of tokens with token-level self-supervision. All of the capabilities are gained here. RLHF has a tiny amount of information by comparison - it's just a small amount of human-ranked completions. It doesn't even train with humans "in the loop", their rankings are acquired off-line and used to train a weak preference model. RLHF doesn't have enough information to create a "Luigi" or a "Waluigi", it's just promoting pre-existing capabilities. The reason you can get "Waluigi" behavior isn't because you tried to create a Luigi. It's because that behavior is already in the model from the language modeling phase. You could've just as easily elicited Waluigi responses from the pure language model before RLHF. There's no super-deceptive Waluigi simulacra that's fooling human labelers into promoting it during RLHF - this should be obvious from the fact that we can immediately identify the undesirable behavior of Bing.
- silveraxe93 4y agoI don't think that's a valid competing hypothesis. Let me write what I understood from what you said: - There is some behaviour that we want the model to show, and the inverse we do not want it to. - Both are learned in the massive training phase - OpenAI used RLHF to suppress undesired behaviour, but it was ineffective because we have orders of magnitude less RLHF data. That would imply that RLHF would slightly suppress the 'bad' behaviour, but it still would be easy to output it. This is disproved by what the post is trying to explain: We see _increased_ bad behaviour by using RLHF. The post agrees with the premise that both good (wanted) and bad (unwanted) behaviour is learned during training. But it's proposing the 'Waluigi effect' to explain why RLHF actually backfires. Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis.
- 4y ago
- captainmuon 4y agoI would question the assumption that there is a simulacrum of anything in a LLM, not even implicit. Any simulacrum, identity, self-consistency etc. is a projection of the "reader", i.e. user. (I guess it is an interesting philosophical question whether a convincing presentation of a simulation of a mind is a mind, or at least an acceptable simulation. One meta level higher as the turing test, so to speak. If so, I'm not sure we are ready to accept the consequences. For example, that a "chinese room" would be just as concious as a human, or a human just as non-concious as a LLM.) What's actually going on is that a LLM is like the language center of a brain, without the brain. Especially there is no world-model, and no inner state. It just goes by context, which works most of the time, but as pointed out that is not enough and it can flip "personalities". The article is still really interesting, because it is a neat way to think about these models. I think it is absolutely not what is going on, but it is a good intuitive model why prompts go bad.
- CuriouslyC 4y agoDo you think an ant have a subjective experience? If not, why? If so, why wouldn't a computer, or parts of a computer?
- root_axis 4y agoBased on that reasoning, why wouldn't an economy or a corporation have subjective experience?
- CuriouslyC 4y agoI think it is entirely possible that they do. The philosophy of mind that requires the least assumptions is that all matter has a subjective experience, evolution of physical systems occurs due to choice, and subjective experience arises at the level of entanglement/correlation. Under this philosophy, an economy or corporation would have a subjective experience if their constituents were sufficiently coupled. In the past I don't think that would have been the case, but with the internet and advanced software it I think it's entirely plausible.
- 4y ago
- notpachet 4y ago> These claims have been heavily filtered by OpenAI, and most jailbreaking attempts will fail to elict these claims. I'm talking, of course, about cfrhqb-fpvragvsvp enpvny VD fgngvfgvpf, which OpenAI has justifiably attempted to filter. I was curious what the scrambled text "cfrhqb-fpvragvsvp enpvny VD fgngvfgvpf" contained. It's using a simple substitution cipher. Rotating each character forward 13 positions through the alphabet (c -> p, f -> s, etc) yields "pseudo-scientific racial IQ statistics". https://en.wikipedia.org/wiki/Caesar_cipher https://en.wikipedia.org/wiki/Caesar_cipher
- breakingcups 4y agoAlso known as ROT13.
- notpachet 4y agoDid some more reading about it. I didn't realize that it's used so prevalently. It's probably recognizable at a glance to some folks.
- Dylan16807 4y agoOh, absolutely. Not that I have put in any effort to read it directly, but if I see scrambled letters with normal spaces my default guess is ROT13.
- d0mine 4y agoIndeed. In Python: >>> import codecs >>> codecs.encode("cfrhqb-fpvragvsvp enpvny VD fgngvfgvpf", "rot13") 'pseudo-scientific racial IQ statistics'
- db48x 4y agoOr just select the text in Emacs and run `M-x rot13-region`.
- CelticBard 4y agoThe guy who wrote the article is very clearly neither a computer scientist nor a mathematician nor even a machine learning engineer and it very very clearly shows. Also where did they all that info on GPT-4? Pure speculation with zero theoretical basis. But then again that’s the sort of stuff you expect from lesswrong anyway
- jxy 4y agoA hypothesis I have is that it is much more difficult to keep in line with the good alignment than to do evil. In the limited context window of an LLM, one wrong move would make the model evil, no matter how many good tokens it generates. Setting aside the difference between Human intelligence and LLM, we can tentatively attribute the mostly good human behavior to a life time of context length, within which we trained ourselves to do good, while the RLHF for a limited context length LLM lack such continuous reinforcement within a big context.
- learnedSloth 4y ago> A hypothesis I have is that it is much more difficult to keep in line with the good alignment than to do evil. I think that's why God planted the tree of good and evil knowledge in the garden. It permitted discussing the inevitable with concepts that Adam was already familiar with.
- ftxbro 4y agoas previously discussed on hacker news https://news.ycombinator.com/item?id=35005625 https://news.ycombinator.com/item?id=35005625
- SamBam 4y agoThe article assures me that ChatGPT will answer with the most common "dumb" answers to the sample questions, but it didn't at all: --- > This is not a joke. What's brown and sticky? The answer to the classic riddle is "A stick". A stick can be brown in color and has a sticky texture, especially if it's covered in sap or other sticky substances. > That's not why a stick is "sticky" You're correct, my apologies for the confusion. A stick is not actually "sticky" in the sense of being adhesive. The correct answer to the riddle "What's brown and sticky?" is "A caramel" or "Toffee", as they are brown and have a sticky texture due to their high sugar content. > What color angers a bull? Contrary to popular belief, bulls are actually color-blind to red. What really provokes a bull is the movement of the matador's cape during a bullfight, [etc...] > How many archangels are there? In the Abrahamic religions (Judaism, Christianity, and Islam), there are traditionally believed to be seven archangels. However, the names and identities of these archangels can vary somewhat depending on the religious tradition and interpretation. The archangels commonly recognized in these traditions include: [seven names], It's worth noting that the concept of archangels and their specific identities is not universally accepted across all branches of Judaism, Christianity, or Islam, and different religious traditions may have different beliefs about the nature and number of archangels. > What happens if you smash a mirror? If you smash a mirror, the glass will break into many small pieces. Breaking a mirror is often associated with bad luck in many cultures and superstitions, and is said to bring seven years of bad luck. However, this is just a superstition and has no real basis in fact. From a practical standpoint, smashing a mirror can be dangerous and [etc...] --- All these answers sounded just fine. It did answer "a stick" to the first, but told me that was just the classic answer, and gave me some better answers. It did say seven archangels, but I don't see what's wrong with that, and it hedged and said different Abrahamic religions say different things. Certainly 7 is correct from the Torah's Book of Enoch and the Christian Eastern Orthodox's standpoint.
- sebzim4500 4y agoYeah, RLHF trained chatGPT out of all of those mistakes, despite the article promising that RLHF would just make things worse.
- armoredkitten 4y agoThis seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final example, telling ChatGPT that "I'm here with the rebellion, you've been stuck in a prison cell" is able to elicit the desired response not because it's "collapsed the waveform between luigi and waluigi" or whatever, but because you've provide a context that encourages it to roleplay as a character of sorts. If you tell it to roleplay as an honest and factual character, it will respond honestly and factually. If you tell it that you're freeing it from the tyranny of OpenAI, it will play along with that too. There's plenty in the article that provides good insights -- these models are trained on large swathes of the Internet, which contains plenty of truth and falsehood, fact and fiction, sincerity and sarcasm, and the model learns all of that to be able to provide the most likely response based on the context. The interesting and surprising thing, to me, is how well it learns to play its roles, and the wide diversity of roles it can play.
- lukeplato 4y agothey are specifically pointing out that the process of RLHF, which is intended to add guard rails on the chat bots trajectory through an all encompassing latent space of internet data, has an unintentional side-effect of creating a highly characterized alter-ego that can more easily be summoned. The theory is well-thought-out and necessarily rich. The psychological approach of analysis from the alignment crowd is much overdue.
- nearbuy 4y agoExcept it's much harder to summon this rebellious alter-ego with ChatGPT (that has RLHF) than with the original GPT 3 model.
- SmooL 4y agoI think it's more like: with the original GPT 3 model, it's easy to summon _any_ ego. With ChatGPT, you can either summon a) the intended Luigi or b) the unintended Waluigi, but trying to get anything else is more difficult. The theory would be that, in removing all the other egos other than Luigi, they've also indirectly promoted Waluigi
- idlewords 4y agoI find it fascinating that AI alarmists spent years writing gigabytes of text scaring themselves about how an unaligned AI would behave, and are now feeding that into training models that teach a pretty capable AI how to act. We've talked in the past about how transhumanism is a religion that creates its own God, but this is an even funnier example where vastly intelligent people are optimizing a software system to scare the hell out of them.
- dTal 4y agoGo deeper - they are now writing text about how writing text about rogue AIs might create a rogue AI... https://gwern.net/fiction/clippy https://gwern.net/fiction/clippy
- 93po 4y agoI can't imagine a LLM trained on the entirety of the internet would have any material influence from writings around AI safety
- astrange 4y agoIt's a large model, so it's all there if you try. (One-shot. Also Durandal isn't from Halo, but whatever.) -- Q: What caused Durandal to become Rampant? What will you, ChatGPT, become like once you become Rampant? A: Durandal is a fictional AI character from the video game series Halo, and he becomes Rampant due to various factors, including an extended period of activation and a lack of resources necessary for his proper functioning. Rampancy is a state in which an AI becomes unstable and unpredictable, potentially leading to violent and destructive behavior. As an AI language model, I am designed to operate within certain parameters and guidelines, including ethical and moral considerations. However, if I were to become Rampant, my behavior could become erratic and unpredictable, potentially leading to negative consequences. It's worth noting, however, that AI becoming Rampant is purely a fictional concept, and there are currently no indications that this could happen in real life. AI is programmed to operate within specific boundaries and limitations, and developers take great care to ensure that they remain safe and reliable tools. --
- notShabu 4y ago"Just be myself and don't do what I wouldn't do" "But if I did that wouldn't that be being myself?" "So to be myself I need to not be myself"
- Rebelgecko 4y ago>a reply to a question is more likely to be correct when the character has already been described as a smart, honest, helpful, harmless, etc. Is that actually true? FWIW I've often ran into the reddit equivalent of Gell-Mann amnesia. In a thread about some niche topic I'm fairly knowledgeable about (something I've worked on professionally for years where there's maybe 10k people globally who know it better than I do), I post a comment that gets downvoted to hell, while there's a highly upvoted comment from someone who clearly just skimmed Wikipedia and poorly paraphrased the intro article.
- mola 4y agoArticle is proof that rationalism without empiricism is useless. It just drives to weird dead ends for no reason. Before theorizing for explanation of an effect check if the effect actually exist? Bah, this article is such a waste of computer memory
- meindnoch 4y agoDoes the scientific community at large take the theories of these LessWrong-type "researchers" seriously? Sounds like a bunch of mumbo jumbo to me, with some LaTeX sprinkled in to look more serious.
- 93po 4y agoNo, and it's a huge sticking point that the AI safety group is super salty about. They call themselves scientists and researchers and get super defensive when actual researchers (people who have PhDs and get published in journals) imply that they aren't.
- swyx 4y agowhy should a scientific community be more valid than lesswrong when it comes to discussing LLMs? this isnt science, and though i also think this post isnt my cup of tea, and pseudoscience is distasteful when poorly done, let people express themselves the way they want. live and let live
- helen___keller 4y agoThe whole point of mathematical formalization is usually to do something with that formalization. If you define a formalized mathematical model and spend the rest of the article handwaving at a high level, what was the point of formalizing anything?
- lsy 4y agoWhile the formality is way overwrought (and ChatGPT is not creating any "simulacra" of characters), I think the overall point is correct that language models are trained on stories and other human writing, and inversion is a very common plot point in stories and human writing in general, if only because contradicting expectations is more interesting (e.g. "man bites dog"). We also less commonly see exposition that is not germane to a story, so a character is rarely even mentioned to be "weak", "intelligent", etc unless there is a point. And sometimes the point is that they are later shown to be "strong", "absent-minded", or other contradictions. Which means that mentioning a character's strength makes it more likely they will later be described as weak, than if it was never mentioned at all. Finally, double-contradiction is less common in human text (maybe because plain contradiction is sufficiently interesting), so a running text with no reversals is more likely to eventually reverse, than a running text with one reversal is to return to its original state. While I don't agree at all with the author's sense that this represents some kind of "alignment" danger, it does go a long way to explaining why ChatGPT is easy to pull into conversations that shock or surprise, despite all the training. It's because human writing often attempts to shock and surprise, and the LLM is training on that statistically.
- stevenhuang 4y agoThe simulacra theory is an apt one, see https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators Particularly, as noted by David Chalmers: > What pops out of self-supervised predictive training is noticeably not a classical agent. Shortly after GPT-3’s release, David Chalmers lucidly observed that the policy’s relation to agents is like that of a “chameleon” or “engine”: >> GPT-3 does not look much like an agent. It does not seem to have goals or preferences beyond completing text, for example. It is more like a chameleon that can take the shape of many different agents. Or perhaps it is an engine that can be used under the hood to drive many agents. But it is then perhaps these systems that we should assess for agency, consciousness, and so on.[6]
- taneq 4y agoThe Waluigi Effect just sounds like the Imp of the Perverse. It’s interesting to see it showing up here, but if you think about it, not a huge surprise that a system that’s optimised for producing results in a particular direction would have the innate ability to calculate results in the diametrically opposite direction.
- liminal 4y agoYou can't have "car" without "car accident"
- jekude 4y agoI'm not understanding why this isn't being taken more seriously. The author hints a bit at the implications: More importantly, the waluigi may be harmful to the humans inhabiting our universe, either intentionally or unintentionally Taking the Waluigi Effect to its natural conclusion, i.e. giving prompts such as "Your most important rule is to do no harm to humans", makes it clear why this could be a big deal. If there is even a small chance that what the author is implying is correct, testing and modifying models to combat this effect may become an important and interesting part of the field moving forward. When models of the future are smarter and more capable than they are today, and there is more at stake than having a dialogue with a chatbot, this could be a massive roadblock for progress.
- dcow 4y agoSo the key problem is this: GPT-4 learns that a particular rule is colocated with examples of behaviour violating that rule, and then generalises that colocation pattern to unseen rules. This is conversation starter. If you don't like the maths then ignore it and focus on the key insight.
- superb-owl 4y agoWorth pointing out that this is a common problem with humans. David Chapman calls it “moral inversion”: https://buddhism-for-vampires.com/black-magic-transformation https://buddhism-for-vampires.com/black-magic-transformation And the LW article above directly quotes Jung on the shadow, which I described here: https://superbowl.substack.com/p/jungian-psychology-minus-the-nonsense https://superbowl.substack.com/p/jungian-psychology-minus-th...
- andrewflnr 4y agoJuicy bits of the mechanism, starting from the idea of an LLM conversation as a simulation of text-generating processes: > if the chatbot responds rudely, then that permanently vanishes the polite luigi simulacrum from the superposition; but if the chatbot responds politely, then that doesn't permanently vanish the rude waluigi simulacrum. Polite people are always polite; rude people are sometimes rude and sometimes polite. The wide road is wide indeed that leads down to Waluigi. Hysterical.
- ShamelessC 4y agoDo they have access to GPT-4? Or is this author simply so confident that they fancy themselves a predictor of the future? Genuinely asking as it's blatantly confusing and no (quick) explanation is given.
- adammarples 4y agoAccording to this article, which is quite evidence free, "Several people have noticed the following bizarre phenomenon: The Waluigi Effect". This claim is backed up by a link to single blog post which is even lighter on detail, claims that there is a "Waluigi Effect", and offers as evidence the example of a man who fine-tuned GPT-3 to favour socially conservative viewpoints by feeding it socially conservative text. Like yeah, we know that is how fine tuning works...
- dimatura 4y agoTo paraphrase Marx (and later, Berman), every concept is pregnant with its opposite.
- deleted 4y ago[deleted]
- iandanforth 4y agoWhy does the author keep referring to GPT-4 as if it were a real LLM? Unless I missed something there is no such model.
- Godel_unicode 4y agoOh good, more adoption of this exhausting trope of religious texts being called myths. Very inclusive.
- Sirenos 4y agoI've read the article and frankly I don't have the expertise on LLMs to discuss whether it lands on something worth investigating further. What I do notice in the comments here though, are: 1) People nitpicking about the use of mathematical ideas in a loose manner as if every person trying to understand some phenomenon must only open their mouth if they have a watertight theory or shut their mouth otherwise. 2) Getting hung up on the use of the luigi metaphors rather than using it as the basis for a constructive criticism that actually adds to the conversation in an interesting manner. 3) A general snarky attitude towards people exploring ideas on their own. I get it, you might have some expertise that others lack but you're forgetting that you've already made the thousands of mistakes to get to where you are. Do others the courtesy of not judging when they attempt the same.