4 ms·
The published transcripts are the most valuable part of this. We've found that real exploit chains almost never look like what you'd dream up internally. One th
by arizza 7mo ago
The published transcripts are the most valuable part of this. We've found that real exploit chains almost never look like what you'd dream up internally. One thing I'd push on is are the agents stateful across attempts? Single-turn exploits are table stakes, but the failures that actually scare me are multi-step sequences where each individual action looks benign and only the session-level pattern is dangerous. That's where prompt-level guardrails completely fall apart and you need enforcement at the action boundary itself.
- zachdotai 7mo agoThe agent isn’t stateful across sessions, but the guardrail layer is — it has access to the full conversation history when evaluating each tool call. So you’d think it would catch exactly the kind of multi-step pattern you’re describing. Have you managed to make it work?
- jamiemallers 7mo ago[dead]