4 ms·
> It is very important to us that the deployment of fine-tuning is safe. To preserve the default model's safety features through the fine-tuning process, fine-
by todd3834 3y ago
> It is very important to us that the deployment of fine-tuning is safe. To preserve the default model's safety features through the fine-tuning process, fine-tuning training data is passed through our Moderation API and a GPT-4 powered moderation system to detect unsafe training data that conflict with our safety standards.
I wish there was some documentation on what kinds of things are determined unsafe. There are plenty of things I think we would all agree are unsafe. I'm sure we don't want fine tuned models on how to cause physical harm on other people.
I don't envy the challenge of making the call for more gray area, sometimes even cultural differences, in what is safe or not. Seems like a very hard problem we've seen social media struggle with. I'm reminded of some of the Covid "misinformation" being deemed as unsafe
- lucasyvas 3y agoI'd like to see this too. I'd hate for AI moderation to become the next generation of "the social media feed algorithm" where it's completely opaque. Trading echo chambers for censorship in that case.
- netruk44 3y agoYou can see the list of things the moderation endpoint scans for in the OpenAI documentation: https://platform.openai.com/docs/guides/moderation/overview https://platform.openai.com/docs/guides/moderation/overview I'm unsure of what the "GPT-4 powered moderation system" entails, though. Conjecture: My unsubstantiated guess would be them prompting GPT-4 with something like "Is the following excerpt considered to be harmful or unsafe: {training data}" and then limiting the output to just a few words like "Yes", "No" and "It's unclear".
- MallocVoidstar 3y agoAlways funny when I see people talk about using LLMs for creative writing when both OpenAI and Anthropic believe that generating any amount of sex or violence is grounds for a ban.