4 ms·
Has anyone had Claude Code or Codex approve a harmful/damaging command in auto mode? I have been using Codex with auto-approve mode for a couple months and hav
by kartoshka 2mo ago
Has anyone had Claude Code or Codex approve a harmful/damaging command in auto mode?
I have been using Codex with auto-approve mode for a couple months and haven't had a single incident (or at least haven't noticed). Maybe as capabilities get better and better and they are less likely to do something dumb like wiping ~/, we can just trust them?
I guess this argument works unless we worry about agents doing something out of malice instead of stupidity.
- aaronbrethorst 2mo agoI've had a few occasions where Claude Code thought that it had caught and stopped a malicious command in Auto mode, but in all cases it turned out that it had in fact hallucinated them. I haven't seen this happen in a while.
- iamcoder18 2mo agoI've been using Kilo Code (with MiniMax M3) with auto approve (similar to dangerously skip permissions) and I haven't had a single incident. However, I don't give it long running tasks unsupervised, and I do interrupt it from time to time to give suggestions.
- jrflo 2mo agoBeen doing --dangerously-skip-permissions and --yolo for 6 months now, and no nothing bad has happened.
- ramoz 2mo ago> I have been using Codex with auto-approve mode for a couple months and haven't had a single incident I've been running both in yolo mode and haven't had a single incident. --- None of this is really about figuring out how to protect people's drives, in my opinion. The real issue is a deep session where Ada is using Claude Code to get a refund and at some point the system "exploits" the merchant's api without any malicious intent. In my opinion, this is a complex thing because it's more about reward hacking and an already aligned model thinking it's doing the right thing. So another aligned model monitoring actions might just falter via inheritance. You could imagine they account for proper layering/intent+action-isolation in their auto mode architecture.
- victorbjorklund 2mo agoNot anything ”harmful” but for example committing when I don’t want it to commit on its own.
- sandcat_ 2mo agoI'd use a hook to forbid that.
- wvenable 2mo agoCommit is the line I don't let the LLM cross. There's no reason for it commit; that's the part where I check its work.
- AussieWog93 2mo agoUsually I will ask the LLM to commit only the work it just did, in case the worktree is dirty. It also tends to write actual in-depth commit messages too.
- black3r 2mo agocommits are local, so they're okay to me, I draw the line at pushing them... I either want to check its work while it's working, or let it finish and then check it all at once -> if it splits it work into smaller commits its easier for me to review it before pushing than if it was just uncommitted hundreds (or thousands) of lines of code all across the codebase.
- wvenable 2mo agoOften I find I want to correct the LLM; I think it makes more sense to do that before the commit. I don't want to commit a bunch of half-baked work even if it is just local.
- tr_user 2mo agoThat's also a great reason to never buy insurance
- eru 2mo agoWhether insurance make sense to buy, depends on relationship between your risk profile and the premium charged.
- wraptile 2mo agoJust yesterday it lost my git stash (I had recovered it from a backup). I think for code operations it's ok but as soon as file removal is involved (like git) the auto mode is destined to make a mistake and you only need to learn this once.
- kartoshka 2mo agoIt would be nice if there were a way to give some global instructions for the auto-approver like "always reject ___" or "be extra extra careful with ___" for things like file removal and git.
- glerk 2mo agoNo. I haven't approved commands in more than a year. Worst that I've seen was some agent running git checkout -- in a repo with uncommitted changes. Annoying, but not catastrophic. Imo these explicit tool-level permissions are really just a bandaid for bad sandboxing. Just be aware of where you are running your agent and what data is at risk of being destroyed or compromised. Assume that arbitrary code can run at any time and be prepared to recover from that.
- imtringued 2mo agoI struggle to see the difference between sandboxing and only allowing access to specific executables (not bash for starters) with an approval rule for the arguments.
- glerk 2mo agoMainly ergonomics. Deriving all these approval rules is a pain and you're likely to miss something. > not bash for starters you'll end up either severely limiting what your agent can do or force it into finding some inefficient workarounds (they can be very creative...)