4 ms·
Yeah, this kind of toxic output sadly still can happen :-/ We have fully analyzed the training dataset (1128 GB) using Detoxify (https://github.com/unitaryai/d
by MasterScrat 5y ago
Yeah, this kind of toxic output sadly still can happen :-/
We have fully analyzed the training dataset (1128 GB) using Detoxify (https://github.com/unitaryai/detoxify https://github.com/unitaryai/detoxify) to filter out problematic content. But of course detecting toxicity is a tough challenge in itself, so this process is imperfect at best.
We are using the RealToxicityPrompt framework (https://realtoxicityprompts.apps.allenai.org/ https://realtoxicityprompts.apps.allenai.org/) to analyse how toxic our models are and to steer our efforts in this direction. This means we are generating thousands of completions and analysing them to see how "nasty" the model is. We plan to write more on this topic soon.
But yeah, this is definitely far from being a solved problem, and our model (as well as all large language models) should be handled with care.
- nud 5y agoCan you share approximately what fraction of the documents got filtered out with your toxicity detection? Also, I wonder what thresholds you used on Detoxify for filtering?