4 ms·
Yeah, this is a real problem that we work to resolve. Partly it's a defence-in-depth approach, and having an LLM check the actions of a coding agent does reduce
by tOOtl 2mo ago
Yeah, this is a real problem that we work to resolve. Partly it's a defence-in-depth approach, and having an LLM check the actions of a coding agent does reduce the likelihood of dangerous actions going through even if it's not perfect. There's also a benefit to using a separate instance of the same model, or a different model that doesn't have correlated failure modes with the agent it's monitoring.
In the cases where you can deterministically block actions, with sandboxes and file permissions, that's better than relying on an LLM. But that doesn't work for all actions, as the OP shows.