3 ms·
Is the solution not sort of easy? You first ask the ai if the input prompt is nefarious with a yes or no question in a first pass. You don't show the user this
by callesgg 4y ago
Is the solution not sort of easy?
You first ask the ai if the input prompt is nefarious with a yes or no question in a first pass. You don't show the user this output.
If the first pass indicates the input prompt is nefarious. You don't continue to the next pass. If the first pass says the input prompt is okay you pass the input prompt to the ai.
I guess it is computationally costly to run the ai twice, but I bet you would get very good results. You might be able to fool the first pass, but then it would be hard to get useful responses in the second pass.
- swyx 4y agodiscussed here https://simonwillison.net/2022/Sep/17/prompt-injection-more-ai/ https://simonwillison.net/2022/Sep/17/prompt-injection-more-... and ive been told theres even more research in academic circles that have been inconclusive