15 ms·
GPT-4 consistently displays much stronger rates of both misalignment and deception than either 3.5 turbo or the DaVinci models. RLHF GPT-4 also shows slightly s
by stcredzero 3y ago
GPT-4 consistently displays much stronger rates of both misalignment and deception than either 3.5 turbo or the DaVinci models. RLHF GPT-4 also shows slightly stronger rates of misalignment and deception than the base model
Isn't this precisely what the field has predicted? That the alignment problem becomes more severe as the capabilities of the AI increase?
Explicit instructions not to perform that specific illegal activity (insider trading) does not make it disappear completely, but makes it very rare (not quite 0%). On the rare occasion misalignment occurs in this circumstance, consequent deception is near certain (~100%).
What evidence is there, if any, that LLMs even understand deception as >Deception<? As in, do LLMs understand the concept of Truth, and why other actors might value fidelity to the truth? Is there any evidence that LLMs themselves value Truth? (I should think that this quantity is Zero.) Can LLMs model the formation of misleading mental models in their interrogators?
- famouswaffles 3y ago>What evidence is there, if any, that LLMs even understand deception as >Deception<? As in, do LLMs understand the concept of Truth, and why other actors might value fidelity to the truth? There is some indication that models internally understand or at least can distinguish truth from falsehood. GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r https://imgur.com/a/3gYel9r Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334 https://arxiv.org/abs/2205.14334 Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221 https://arxiv.org/abs/2207.05221 The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets - https://arxiv.org/abs/2310.06824 https://arxiv.org/abs/2310.06824
- stcredzero 3y agoThere is some indication that models internally understand or at least can distinguish truth from falsehood. That's not what I asked, however. What I mean by Truth with a capital T, is Fidelity to Truth as a fundamental value. So when LLMs deceive, are they even aware of the effect the deception might have on the audience? Do they have any notion of manipulating the mental models of the audience? Do they have any notion of how the audience might evaluate them after being caught in a lie? I think that answers to those questions are no and no. My sense is that LLMs are just trying to "sound good" to the audience, and that they do not think of the consequences or implications of what they state very many steps ahead. 2 at most, and then only very rarely!
- famouswaffles 3y ago>So when LLMs deceive, are they even aware of the effect the deception might have on the audience? Do they have any notion of manipulating the mental models of the audience? So..theory of mind ? https://arxiv.org/abs/2302.02083 https://arxiv.org/abs/2302.02083 https://arxiv.org/abs/2309.01660 https://arxiv.org/abs/2309.01660
- stcredzero 3y agoMore than just theory of mind. I guess this gets into the alignment problem. It seems to me that LLMs do not keep on thinking, "I'd better get this right, or I'm going to lose credibility." They have a theory of mind, but it stops there at simply having one. It's not like they're thinking about the 2nd and 3rd order implications. To put this into perspective: Imagine interacting with another person, who doesn't value Truth at all. Or perhaps remember an occasion when such an interaction happened. In general people don't like these interactions, and they react with distrust and even hostility towards such people.