3 ms·
Do open weight models have similar content gaurdrails in place?
by DontchaKnowit 5mo ago
Do open weight models have similar content gaurdrails in place?
- yk 5mo agoNo, but actually yes. Guardrails usually refers to a step in the inference pipeline where you check that it is consistent with policy while open weight models don't come with such a multistep pipeline. However open weight models are aligned during RLHF step, which means they will refuse to discuss overly sensitive topics. There are techniques to remove those, if you look for uncensored models on huggingface.
- benkaiser 5mo agoOften there are "abliterated" or "uncensored" tuned models that suppress the rejections. From my high level understanding it is performed by finding which weights activate for the rejection and lowering those so the model is less likely to reject. It doesn't fix if the model doesn't know what you're asking it though (i.e. if the model never actually learned about meth production in the first place).
- ndr_ 5mo agoYes. OpenAI's GPT-OSS was training using Deliberative Alignment (which was found to be flawed in a competition on Kaggle, but still). https://arxiv.org/abs/2412.16339 https://arxiv.org/abs/2412.16339