3 ms·
Besides the debate about whether this is "safer" than manual human review, I have a slightly different problem. Very often, when I'm running Claude in manual r
by dgunay 2mo ago
Besides the debate about whether this is "safer" than manual human review, I have a slightly different problem.
Very often, when I'm running Claude in manual review mode, it will attempt to do things which are not "dangerous" but are misaligned with what I want it to do. Maybe I'm fighting the model here but for example, when orchestrating other agents to do work, Claude really badly wants to be overly prescriptive about how the work gets done, telling them exactly which files to edit, exactly what not to do, etc. instead of trusting the guardrails, review agents, or humans in the process to catch code-level mistakes. And no, telling it not to do this does not stick. Manual review is the last line of defense I have here.
I have stuff I don't want blacklisted, only allow it to use tools with limited ability to boss around agents, and various hooks to try and catch behavior that the permissioning system can't. If I use Auto mode though, I lose this control. The classifier will gleefully approve these types of commands because guess what, it's also Claude.
- jaggederest 2mo ago> Maybe I'm fighting the model here > And no, telling it not to do this does not stick. You're fighting the model, don't argue with city hall. Set the standards and let it figure out how to execute, stop getting bogged down in the minutia. I try, as much as I can, to treat the session as a black box - only the inputs and outputs matter, internal prompting of subagents is way out of scope. You can't change it via prompt, and you can't control the guardrails, so something else has to give - either your perspective or the system you're managing. If you really believe that the internal prompting is bad, turn off subagents and workflows and only let it execute in thread. But if you're going to do that, you'd better benchmark it against not doing that, because historically fighting the harness and model globally makes everything worse. I would bet you that the subagent prompting is excellent, and anything you do to change it will make it worse, but I wouldn't make it a large bet.
- user43928 2mo agoI think he has a point. I noticed the same with skills that invoke another agent harness. I just have a skill to review the changes in the current worktree. By default, it will put lots of instructions about locating the changes into the prompt, like explaining how to use git diff. These instructions are obviously unnecessary. I can believe that the same issue of needlessly verbose prompts might exist with subagent spawning. I would not go to customize that one however. With skills, it is a more natural fix.
- drdec 2mo agoIMHO, the verbosity is an attempt to make the results more deterministic. The more specific the prompt, the less wiggle room.
- jaggederest 2mo agoRemember, it's not just about the model "knowing" something. It's about the frame of mind it's in, if I may extend a metaphor. Including git diff in the prompt doesn't just teach it about git diff, it makes it more likely to use it. A lot of this is obsolete with newer models, but funnily enough, the newer models don't know that - their theory of mind is at least partially obsolete in itself, so they continue to do the old "you are a software architect, don't make mistakes" prompt style.
- adrian17 2mo agoIf I ask a model to do a change involving editing a file, and it starts investigating internals of my build system, then sure it might not be counterproductive or break the task, and might have taken only extra 30 seconds; but for all I know, my quick rejection of a shell invocation (with a simple "irrelevant to the task" comment) might have just saved me half of today's Opus tokens, which already makes it worth it.
- dannyw 2mo agoYou can use another harness like Pi, OpenCode, etc and build your own auto reviewer if you’d like (or adapt the open source Codex one). If you have an openai subscription you are explicitly allowed to use your subsidised tokens / usage limits with any harness you like, not just Codex. Unfortunately this is a violation of Anthropic’s terms but that’s their business decision.
- dan_t 2mo ago[dead]
- TZubiri 2mo agoControl is not only about security, you don't review the work of your employees just because you want to avoid them stealing from the register, you want to perform QA on their tasks and ensure they are aligned.
- thinkingtoilet 2mo agoThere is no debate. It is not safer. I have to worry about HIPAA compliance and despite cluad files and some restrictions in the settings.json it still will try to violate HIPAA compliance from time to time. I literally can't make a single mistake. If you leave it on auto, it will mess up eventually.
- brad-mcevilly 2mo ago[flagged]