3 ms·
Yeah I agree it doesn't show increased bad behaviour. It's definitely a weak point in the argument. The theory is extremely interesting though. And better yet,
by silveraxe93 4y ago
Yeah I agree it doesn't show increased bad behaviour. It's definitely a weak point in the argument.
The theory is extremely interesting though. And better yet, it's falsifiable!
If someone went around compared an RLHF model vs non-RLHF and found them equally likely to 'Waluigi' then we'd know this is false. And conversely if we found the RLHF more likely to Waluigi then it's evidence in favour.
The asymmetry in the hypothesis is really nice too.
If this was true then I'd expect it to be possible to flip the sign in the RLHF step, effectively training it in favour of 'bad' behaviour. Then forcefully inducing 'Waluigi collapse' before opening to the public!
- skybrian 4y ago"Flipping the sign" implies the existence of an internal representation that we can't know about from the outside. Since all we see are the words, I prefer to call it a plot twist. Language models are trained on a large subset of the Internet. These documents contain many stories with many kinds of plot twists, and therefore it makes sense that a large language model could learn to imitate plot twists... somehow. It would be interesting to know if some kinds of RLHF training make it more likely that there will be certain kinds of plot twists. But there are more basic questions. What do large language models know about people, whether they are authors or fictional characters? They can imitate lots of writing styles, but how are these writing styles represented?
- taneq 4y agoI propose we take this further and adopt this phrasing for all unanticipated software behaviour. ATM says you have (uint32_t)-403 cents in your account? Plot twist. Self driving car pathing road-runner style through a billboard of a tunnel? Plot twist!