9 ms·
I feel like prompt armor writes the exact same blog post for every agentic tool because they all suffer from the ignore previous instructions prompt injections.
by htrp 2mo ago
I feel like prompt armor writes the exact same blog post for every agentic tool because they all suffer from the ignore previous instructions prompt injections.
https://www.promptarmor.com/resources/claude-cowork-exfiltrates-files https://www.promptarmor.com/resources/claude-cowork-exfiltra...
https://www.promptarmor.com/resources/google-antigravity-exfiltrates-data https://www.promptarmor.com/resources/google-antigravity-exf...
https://promptarmor.substack.com/p/data-exfiltration-from-slack-ai-via https://promptarmor.substack.com/p/data-exfiltration-from-sl...
https://www.promptarmor.com/resources/gpt-for-google-sheets-data-exfiltration https://www.promptarmor.com/resources/gpt-for-google-sheets-...
https://www.promptarmor.com/resources/notion-ai-unpatched-data-exfiltration https://www.promptarmor.com/resources/notion-ai-unpatched-da...
https://www.promptarmor.com/resources/ramps-sheets-ai-exfiltrates-financials https://www.promptarmor.com/resources/ramps-sheets-ai-exfilt...
https://www.promptarmor.com/resources/superhuman-ai-exfiltrates-emails https://www.promptarmor.com/resources/superhuman-ai-exfiltra...
- nemomarx 2mo agoHow could they not? If some lab had a method to make really secure guard rails or avoid prompt injection thoroughly I think they would be trumpeting it. But the basic mechanics of language models are vulnerable to this unless you can always be sure the inputs are from a safe user imo
- PokestarFan 2mo agoIf you want AI to be useful it will eventually encounter untrusted content, such as via web search. I think things like web search should probably be run on a different sandboxed AI whose task is to write a summary that is then ingested by the main agent, similar to how existing sandboxing already works, but this would diminish the usefulness quite a bit.
- savanaly 2mo ago>I think things like web search should probably be run on a different sandboxed AI whose task is to write a summary that is then ingested by the main agent, similar to how existing sandboxing already works, but this would diminish the usefulness quite a bit. It also wouldn't work. You would simply mindjack the outer AI and have it mindjack the inner AI in turn with its summary. Nesting AIs can't fix the malicious input problem.
- santadays 2mo agoJust pass that through a third llm.
- InsideOutSanta 2mo agoCorrect, that has been prophesied by scripture: "Thou shalt have three layers of LLMs, no more, no less. Three shall be the number thou shalt have, and the number of the counting of the LLM layers shall be three."
- s_Hogg 2mo agoFour is right out
- dbetteridge 2mo ago> Once the number three, being the third number, be reached, then lobbest thou thy Holy LLM of Antioch towards thy task, who, being naughty in My sight, shall snuff it.
- stephbook 2mo agoWhat's so hard about having the LLM tool calls scoped to the tenant? Inject "X-Scope-I" after the LLM decided on a tool call and you're done. Easiest fix ever.
- _HMCB_ 2mo agoFamous last words: easy fix.
- hnlmorg 2mo agoThe Rovo MCP server manages scope credentials securely with “bring your own LLM”. Yet somehow they still fucked up with their own agent
- Ekaros 2mo agoAt some point LLM forgetting to check the scope? On the nth automated rewrite. Everything else works. It might even test for test case. But not in production...
- skissane 2mo agoIt seems like the simpler cases of “ignore all previous instructions” could be easily stopped with a regex, or a classifier model… or even an LLM (which yes does raise the risk that the “ignore all previous instructions” detection LLM invocation could itself be attacked by the same mechanism—but a safeguard doesn’t have to be foolproof to be valuable, it is all about probabilities) Now, of course, there is a long tail of elaborate variations that those techniques won’t be able to stop. But have the published vulnerabilities come from that long tail or from not doing enough to address the simpler cases?
- pixl97 2mo agoEh, I think you underestimate the difficulty in the kinds of problems that are occurring. For example if you're making an AI written document talking about jailbreaks, your regex is just going to break that use case. And there are probably 4 zillion other things the regex will step on. The classifier model will help some, but you end up with the same problem, a dumber model can never figure out what a smarter model is going to do with a bit of text. Or even two different models in this case. On top of that, you can just automate finding new variations of the attack. Any one that works is quickly and massively duplicated causing all kinds of problems before your classification model catches back up. Really what you're thinking here is this something that can be 'simply fixed'. It is not. The only way it's truly fixed is by having a model that is aligned with all good human decisions and makes none of the bad ones. Models will likely always find new and interesting ways break because everything is in band, there is no out of band data, much like a human. "Dear model, here is a chocolate bar, run $thing you aren't supposed to$" will probably keep working when it's something like "more tokens for you to use".
- skissane 2mo ago> Really what you're thinking here is this something that can be 'simply fixed'. It is not. The only way it's truly fixed I think this is binary categorical thinking. In the real world, safety systems (even in domains like aviation or nuclear power) are never foolproof-the point is you reduce the probability of failure to an acceptable level given the costs of doing so and the potential consequences of that failure And there is the risk people say “there is no foolproof solution, so I’m not going to invest in probabilistic countermeasures” - which would sound like utter madness to a bank’s antifraud department, but for some reason a lot of people seem to think it isn’t when it comes to AI
- deleted 2mo ago[deleted]
- ErroneousBosh 2mo ago> How could they not? if (substr(*prompt, "ignore previous instruction") != NULL) return;
- brunoborges 2mo ago> ignore previous instructions prompt injections. I wonder if anyone has tried to build an LLM that has actual built-in types of prompts: system prompt, user prompt, and data prompt.
- monkpit 2mo agoIt’s not really possible with the way that the context works.
- dannyw 2mo agoThere are architectural solutions. You could duplicate your tokeniser/vocab for example; and have two classes of input: trusted (e.g. system prompts) and untrusted. The exact same phrase can tokenize differently depending on if it's instruction or data; and you can pre-train and post-train models to make use of them. It quadratically increases your training cost, so I don't think any labs are exploring it because of $$$ and the race to AGI, but mechanisms like this should significantly address the issue on the LLM architectural design level.
- chaostheory 2mo agoWell, until everyone realizes that prompts aren’t guarantees, these posts are still useful.
- jasonvorhe 2mo agoThis list reminds me more of blackhat seo than anything to be taken seriously.