2 ms·
My team is working on Watcher to deal with exactly this. We know that Claude is occasionally going to do stuff we really don't want, but at a rate that's way to
by tOOtl 2mo ago
My team is working on Watcher to deal with exactly this. We know that Claude is occasionally going to do stuff we really don't want, but at a rate that's way too low for manual approvals to make sense, so we hook into Claude Code (or Codex) to approve commands in a way that's a lot closer to `--dangerously-skip-permissions` but without the danger. We use a hierarchy of deterministic rules and heavily-tested LLM monitors to balance speed, cost, and accuracy.
https://watcher.apolloresearch.ai/ https://watcher.apolloresearch.ai/
- tsimionescu 2mo ago"We don't trust the llm, so we built a tool that uses the llm to check if the llm can be trusted"
- tOOtl 2mo agoYeah, this is a real problem that we work to resolve. Partly it's a defence-in-depth approach, and having an LLM check the actions of a coding agent does reduce the likelihood of dangerous actions going through even if it's not perfect. There's also a benefit to using a separate instance of the same model, or a different model that doesn't have correlated failure modes with the agent it's monitoring. In the cases where you can deterministically block actions, with sandboxes and file permissions, that's better than relying on an LLM. But that doesn't work for all actions, as the OP shows.