3 ms·
For little effort the defending dev can ask a second layer of LLM to check if an output is explicitly toxic, filter it, and nullify the red team. Filtering shi
by courseofaction 3y ago
For little effort the defending dev can ask a second layer of LLM to check if an output is explicitly toxic, filter it, and nullify the red team.
Filtering shitty content is easier than creating it with a properly constructed LLM system, the complaints about toxic outputs seem to me to be analogous to an electrical engineer complaining that the voltage from the mains is wrong for their device, but refusing to google what an (electrical) transformer is.
Toxic writing pre-exists LLMs. LLMs output writing. This is not a new problem, but we have a new solution - LLM filtering.
- Timon3 3y ago> Filtering shitty content is easier than creating it with a properly constructed LLM system Can you explain this? It feels completely wrong - even OpenAI, who probably have invested the most, can't filter out all "shitty content". An LLM can on the other hand create "shitty content" incredibly easy - even if the creators try to stop it! So how is filtering easier than creating?