3 ms·
the site makes it pretty clear in multiple places that they're talking about "added" or "hallucinated" toxicity. maybe your culture war outrage is misplaced?
by beardicus 3y ago
the site makes it pretty clear in multiple places that they're talking about "added" or "hallucinated" toxicity. maybe your culture war outrage is misplaced?
- Domenic_S 3y agoOk so I know nothing about how this works. It seems like if the model was able to properly detect words in the first place, it would never hallucinate 'toxicity'; if it can't recognize the word with high probability, how will it know whether the speaker actually said $toxicWord or whether it should print something else? Perhaps it's taking a Big List of Naughty Words and weighting them so that the system must be "extra sure" that's what the speaker said, or else fall back to a G-rated word?
- numpad0 3y agoMaybe it's for preventing unwarranted fucks[1]? Translation is more than just concatenating dictionary definitions, and machine translations routinely make this kind of out-of-place and technically correct lookups. 1: https://www.google.com/search?q=engrish+fucking+sign&tbm=isch https://www.google.com/search?q=engrish+fucking+sign&tbm=isc...
- mortimerp9 3y agoMeta employee here. The system is not perfect, or it would not "hallucinate", while it's pretty good, it does sometime make errors (not just hallucination, maybe just some mistranslation due to noise in the training data). What we want is to avoid these errors to introduce toxicity (think swear words) that weren't in the input as this could be very bad for the user. There is a separate system that double checks the output (compared to the input) and tells the translation model to try again if it's too bad.