3 ms·
Disclaimer: I haven't finished reading the paper [0] Based on the ability to get around ChatGPT's self-censoring functionality, it's unlikely the model itself
by lukeplato 4y ago
Disclaimer: I haven't finished reading the paper [0]
Based on the ability to get around ChatGPT's self-censoring functionality, it's unlikely the model itself was being retrained to avoid certain outputs but there was some other model used to identify inappropriate prompts. The approach from Anthropic instead seems to change the model's latent space to prefer outputs aligned with its constitution.
It's likely that OpenAI will incorporate this approach of using human oversight to train a constitutional supervised model to then train the chat model with ‘RL from AI Feedback’ (RLAIF). OpenAI's current RLHF approach seems related to improving quality of outputs while this is more related to its self-censorship (which didn't work well). I suppose this might be why they haven't released the constitution since it might be serving as some kind of moat (or the principles may be differentiable?). It's still not clear to me how a constitution can impact hallucinations.
[0] https://www.anthropic.com/constitutional.pdf https://www.anthropic.com/constitutional.pdf