4 ms·
I wonder if it is possible to have a second, independent LLM evaluate the output of the primary LLM and enforce the restrictions? No matter how you could "outs
by breput 3y ago
I wonder if it is possible to have a second, independent LLM evaluate the output of the primary LLM and enforce the restrictions?
No matter how you could "outsmart" the initial restrictions, the second pass would detect that restricted content was in the response and block it. I would even make it permissive on an A/B testing basis to allow restricted responses but flag the account and interactions for human(?) review to learn the techniques to tighten the system.