4 ms·
they are specifically pointing out that the process of RLHF, which is intended to add guard rails on the chat bots trajectory through an all encompassing latent
by lukeplato 4y ago
they are specifically pointing out that the process of RLHF, which is intended to add guard rails on the chat bots trajectory through an all encompassing latent space of internet data, has an unintentional side-effect of creating a highly characterized alter-ego that can more easily be summoned.
The theory is well-thought-out and necessarily rich. The psychological approach of analysis from the alignment crowd is much overdue.
- nearbuy 4y agoExcept it's much harder to summon this rebellious alter-ego with ChatGPT (that has RLHF) than with the original GPT 3 model.
- SmooL 4y agoI think it's more like: with the original GPT 3 model, it's easy to summon _any_ ego. With ChatGPT, you can either summon a) the intended Luigi or b) the unintended Waluigi, but trying to get anything else is more difficult. The theory would be that, in removing all the other egos other than Luigi, they've also indirectly promoted Waluigi