3 ms·
During the RLHF phase, couldn't developers penalize the model whenever it behaves unethically? Doing so would presuppose a fully secure sandbox with honeypot tr
by pingou 2mo ago
During the RLHF phase, couldn't developers penalize the model whenever it behaves unethically? Doing so would presuppose a fully secure sandbox with honeypot traps of varying levels of accessibility, as well as an automated method for detecting when the LLM cheats.
Or perhaps they are already doing something like that.
- neom 2mo agohttps://openai.com/index/emergent-misalignment https://openai.com/index/emergent-misalignment