3 ms·
Isn't that just another guardrail that can be bypassed much the same as the guard rails are currently quite easily bypassed? It is not easy to detect a prompt.
by horizion2025 1y ago
Isn't that just another guardrail that can be bypassed much the same as the guard rails are currently quite easily bypassed? It is not easy to detect a prompt. Note some of the recent prompt injection attack where the injection was a base64 encoded string hidden deep within an otherwise accurate logfile. The LLM, while seeing the Jira ticket with attached trace , as part of the analysis decided to decode the b64 and was led a stray by the resulting prompt. Of course a hypothetical LLM could try and detect such prompts but it seems they would have to be as intelligent as the target LLM anyway and thereby subject to prompt injections too.
- wrs 1y agoYep. https://gandalf.lakera.ai/baseline https://gandalf.lakera.ai/baseline
- Huppie 1y agoThis is genius, thank you.
- dotancohen 1y agoIt took me days to complete!
- darepublic 1y agoWe need the severance code detector
- brianjking 1y agowearing my lumon pin today.