4 ms·
This is heavily sensationalized. They trained a model to be deceptive, alignment techniques used didn't remove the deception. It's a valuable experiment, but n
by notnullorvoid 3y ago
This is heavily sensationalized.
They trained a model to be deceptive, alignment techniques used didn't remove the deception. It's a valuable experiment, but not that surprising.
- RoboTeddy 3y agoIt was certainly an unresolved question before they did this work! Naively, it seems reasonable to believe that if you adjust all the weights of a neural net towards the behavior you want via SFT and RLHF, that it would compete with/mute/obscure undesired behavior like a back door. But it seems not to be so… Indeed the cute mask does not cover the entire shoggoth— it may still have tentacles (https://images.app.goo.gl/YW9g3BvwGqGwYTgd6 https://images.app.goo.gl/YW9g3BvwGqGwYTgd6)