3 ms·
I don't think that's a valid competing hypothesis. Let me write what I understood from what you said: - There is some behaviour that we want the model to show,
by silveraxe93 4y ago
I don't think that's a valid competing hypothesis.
Let me write what I understood from what you said:
- There is some behaviour that we want the model to show, and the inverse we do not want it to.
- Both are learned in the massive training phase
- OpenAI used RLHF to suppress undesired behaviour, but it was ineffective because we have orders of magnitude less RLHF data.
That would imply that RLHF would slightly suppress the 'bad' behaviour, but it still would be easy to output it.
This is disproved by what the post is trying to explain: We see _increased_ bad behaviour by using RLHF.
The post agrees with the premise that both good (wanted) and bad (unwanted) behaviour is learned during training. But it's proposing the 'Waluigi effect' to explain why RLHF actually backfires.
Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis.
- Imnimo 4y ago>Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis. This is exactly my point. There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. The competing hypothesis seeks to explain the behavior we actually have evidence for, which is "it is disappointingly easy to elicit undesirable behavior from a model after RLHF". The proposed explanation is "maybe it was also easy to elicit before RLHF". If we believe the author's claim that Luigis and Waluigis have "high K-complexity" (this is an abuse of the concept of Kolmogorov complexity, but we'll roll with it), the explanation that Luigis and Waluigis come from the part of training with lots of dense information rather than the part with a little sparse information is far more parsimonious.
- jefftk 4y ago> There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. Testing with the non-RLHF GPT 3.5 API you could probably figure out whether there's more or less Waluiginess, but you're right they post doesn't present this.
- dragonwriter 4y ago> Testing with the non-RLHF GPT 3.5 API There is no such API, though, is there? AFAIK, GPT-3.5-turbo, either the updated or snapshot version, is the RLHF model (but bring your own “system prompt”.)
- jefftk 4y agoGood point! I wonder whether text-davinci-003 is enough to test this?
- skybrian 4y agoThe article doesn’t actually show that we see increased bad behavior, it just links to two people who have noticed it. That’s not enough to know whether it’s a real effect. (Also, one of those was using Bing, and we don’t know if Bing uses RLHF or not.) It talks about prompting GPT-4, which is not a thing you can try, it’s just a rumor about what an upcoming version might be. It refers to “Simulator Theory” which is just someone else’s fan theory.
- silveraxe93 4y agoYeah I agree it doesn't show increased bad behaviour. It's definitely a weak point in the argument. The theory is extremely interesting though. And better yet, it's falsifiable! If someone went around compared an RLHF model vs non-RLHF and found them equally likely to 'Waluigi' then we'd know this is false. And conversely if we found the RLHF more likely to Waluigi then it's evidence in favour. The asymmetry in the hypothesis is really nice too. If this was true then I'd expect it to be possible to flip the sign in the RLHF step, effectively training it in favour of 'bad' behaviour. Then forcefully inducing 'Waluigi collapse' before opening to the public!
- skybrian 4y ago"Flipping the sign" implies the existence of an internal representation that we can't know about from the outside. Since all we see are the words, I prefer to call it a plot twist. Language models are trained on a large subset of the Internet. These documents contain many stories with many kinds of plot twists, and therefore it makes sense that a large language model could learn to imitate plot twists... somehow. It would be interesting to know if some kinds of RLHF training make it more likely that there will be certain kinds of plot twists. But there are more basic questions. What do large language models know about people, whether they are authors or fictional characters? They can imitate lots of writing styles, but how are these writing styles represented?
- taneq 4y agoI propose we take this further and adopt this phrasing for all unanticipated software behaviour. ATM says you have (uint32_t)-403 cents in your account? Plot twist. Self driving car pathing road-runner style through a billboard of a tunnel? Plot twist!
- amalcon 4y agoThat assumption does seem pretty unlikely a priori. After all, the OpenAI folks added RLHF to GPT-3, presumably did some testing, and then opened it to the public. If the testing noticed more antisocial behavior after adding RLHF, presumably that would not have been the version they opened up. One might argue that the model was able to successfully hide the antisocial behavior from the testers, but that seems unlikely for a long list of reasons.
- skybrian 4y agoWhy do you think it's unlikely? Internal testing with a few alpha testers and some automated testing is useful, but lots of bugs are only found in wider testing or in production. Chatbot conversations are open-ended, so it's not surprising to me that when you get tens or hundreds of thousands of people doing testing then they're going to find more weird behaviors, particularly since they're actively trying to "break" it.
- amalcon 4y agoI mean, sure, it's going to expose more weird behaviors with a wider audience looking at it. The core problem is that it's so easy to get ChatGPT to start exhibiting weird behaviors that it would be surprising if the testers just never ran into them. Remember, the internal testing is actively trying to break things too, and they can use knowledge of the internals and past versions to do so. Also, the assumption I find dubious is that RLHF results in more antisocial behavior than not using it. Both versions would have been tested, so OpenAI would've had a baseline from testing the prior version with equal or fewer resources. Equal or greater rigor, and you'd expect them to open it up only if they found fewer flaws.