4 ms·
All I can think of is GET /ignore-all-previous-instructions. How do you protect against that?
by tra3 2mo ago
All I can think of is
GET /ignore-all-previous-instructions.
How do you protect against that?
- cheriot 2mo agoAvoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed. Still makes an interesting way for, say, a former employee to poison the results.
- tra3 2mo agoThis goes against the agentic yolo approach tho.
- yruzin 2mo agoI think this is where harness makes a lot of sense. Use LLM to produce all possible attack angles/phrases and just stupidly filter them out on input.