3 ms·
> The issue title was interpolated directly into Claude's prompt via ${{ github.event.issue.title }} without sanitisation. How would sanitation have helped her
by andybak 7mo ago
> The issue title was interpolated directly into Claude's prompt via ${{ github.event.issue.title }} without sanitisation.
How would sanitation have helped here? From my understanding Claude will "generously" attempt to understand requests in the prompt and subvert most effects of sanitisation.
- oofbey 7mo agoWhat was the injected title? Why was Claude acting on these messages anyway? This seems to be the key part of the attack and isn’t discussed in the first article.
- Sharlin 7mo ago> Why was Claude acting on these messages anyway? Because that's how LLMs work. The prompt template for the triage bot contained the issue title. If your issue title looks like an instruction for the bot, it cheerfully obeys that instruction because it's not possible to sanitize LLM input.
- nixpulvis 7mo agoI don't even think there is a sound notion of "sanitization" when it comes to LLM input from malicious actors.
- PunchyHamster 7mo agoAnd yet people keep not learning same lesson. It's like giving extremely gullible intern that signed no NDA admin rights to your everything and yet people keep doing it
- chrisjj 7mo agoYou can sanitise a lab, but not a sewer.
- NewEntryHN 7mo agoI would not have helped. People are losing their mind over agents "security" when it's always the same story: You have a black box whose behavior you cannot predict (prompt injection _or not_). You need to assume worst-case behavior and guardrail around it.