3 ms·
What are some good prevention mechanisms for this? A sort of firewall for prompts? I've seen people recommend LLMs, but that seems like it wouldn't work well. W
by mosselman 1y ago
What are some good prevention mechanisms for this? A sort of firewall for prompts? I've seen people recommend LLMs, but that seems like it wouldn't work well. What is the industry standard? Or what looks promising at least?
- hoppp 1y agoNothing yet. Probably a new kind of model needs to be trained that can find injected prompts, sort if like an immune system for LLMs. Then the sanitized data can be passed to the LLM after. No real solution for it yet. I would be interested to try to train a model for this but no budget atm.
- yencabulator 1y agohttps://simonwillison.net/tags/lethal-trifecta/ https://simonwillison.net/tags/lethal-trifecta/
- m-hodges 1y agoI have bad news https://matthodges.com/posts/2025-08-26-music-to-break-models-by/ https://matthodges.com/posts/2025-08-26-music-to-break-model...