6 ms·
Breaking Claude Code Opus 5 Auto Mode
- julien_dev 1mo agoI'm quite surprised that we are not seeing something like this more in the wild. Quite concerning
- mkurz 1mo agoMaybe it is used in the wild, but we just don't know.
- rcxdude 1mo agoIt's not that far off a typical trojan, just one tailored to Claude's habits. A lot of the same limits apply.
- chmod775 1mo agoWhat vibe "don't even gotta read it" coder would even notice if it happened?
- 1123581321 1mo agoUsually auto mode does mild things that are dismaying but not dangerous, or maybe not noticed (like if it posts sensitive local data in a request that gets logged but never exploited.)
- comboy 1mo agoInteresting attack, very nicely designed. Not sure if it's much related to the auto mode itself though.
- kevsim 1mo agoThe point is that auto mode gives people a false sense of security that leads them to believe they don't need to run Claude in a proper sandbox. This same attack running in a sandbox (even in YOLO mode) would be comparatively harmless.
- bombcar 1mo agoAuto mode is for people who just keep hitting "YES" on everything, it's a bit better than that. But it's real easy to give auto mode instructions (like "always ask before deploy") and then bypass that just normally.
- Xunjin 1mo agoUntil the model updates or you switch between them often that stops obeying your commands and you have to remind it. In one of the occasions it opened a bug report for me just waiting for hit the enter button.
- whstl 1mo agoI'm not sure I agree. It's easy to "give" instructions, but Claude routinely "forgets" to follow certain instructions, such as "always using the Edit Tool". Just this week it started to use bash with string concatenation to work around some commands that were blocked in settings.json
- silversmith 1mo agoWhat seems to work for me is automation - read file hook that re-injects instructions in the prompt every 15 minutes. Switch on the filename and get language-specific instructions too.
- whstl 1mo agoI have something that injects my relatively small prompt every message, and it still disobeys me after 10 messages or so. The violation above was precisely in this situation :/
- ealready_value 1mo agoEver since they made auto-mode default I swear claude has tuned to use python commands instead of the Edit Tool to frustrate the ~security conscience~ luddites into using auto-mode.
- yeputons 1mo agoI don’t think it’s related to the auto mode at all. It would work perfectly in the manual mode. It does not even need Claude: just give a human a similar archive and hope they run some simple Python from the directory at least once. And make sure there are lots of files do they don’t notice a weird .py around
- nasretdinov 1mo agoThat's an interesting technique! I'd also like to point out that there's something odd with the page itself too, my phone got really hot while I was reading the page, and drained a significant amount of battery charge as well.
- Phemist 1mo agoThis default-to-auto-mode and the misleading marketing is begging for a class action once damages accumulate. Especially considering the Auto Mode even can actively prevent the clean-up!
- thewhitetulip 1mo agoWell, the jokes on us because laws don't apply to AI firms
- olejorgenb 1mo agoIIRC you give up the right to form class action suits by accepting the terms of use?
- bewareofscams 1mo ago> Boris Cherny from Anthropic recently posted that layered defenses could reduce indirect prompt injection on unseen attacks to approximately zero. > I got attack success rates up to 80% using a small sample size. Snake oil salesman misrepresents the data. Color me surprised! /s
- deleted 1mo ago[deleted]
- hahn-kev 1mo agoAs a non Python dev this seems like very surprising behavior for a system library to be modified by just having a file with a specific name in the same folder.
- rcxdude 1mo agoYou would get a similar thing in C and C++ with a system header in a library directory (maybe some compilers would warn on such a thing?). Most languages don't privilege their standard libraries in a way that would prevent this.
- mostlylurks 1mo agoPrivileging the standard library would be a pretty bad way of fixing the issue. It wouldn't prevent the same issue from affecting non-standard libraries, which are also subject to the same issue. The correct way to avoid this issue would be to require local code to be imported in a distinct manner from installed libraries, with an explicitly defined relative path, which is how it works in the javascript ecosystem. If you want to import local code, you just `import foo from './foo';` (for a module in the same folder, `import foo from '../foo';` for module from parent folder, etc.), and if you want to import an installed library, `import foo from 'foo';`.
- rcxdude 1mo agoYeah, I wasn't recommending it as a mitigation per se. python /does/ have relative imports but they're not required. C and C++ notionally have a similar thing with "include.h" vs <include.h> but the behavior is complex and not really designed for avoiding confusion about where the header is coming from.
- MereInterest 1mo agoIt's an awful problem, and a pretty big gotcha for the ecosystem. For example, I worked on a tool that used `tool_name/random.py` for analysis of RNG usage. However, any command run within that directory would shadow the builtin `import random`. As a result, I couldn't run the `black` formatter from within that directory, because it would accidentally import the local `random.py`. The solution is to update the python flags to include `-I`, so that python will run in an isolated mode. <rant> This is relates to my frustration with how PEP-668 was implemented. Python has long had a problem with accidental overwriting of system libraries. If you `sudo python -mpip install foo`, then that can interact very poorly with your distro's `sudo apt-get install python-foo`, since pip would add/remove files that were expected to be managed solely through `apt-get`. But in adding a warning to prevent this, they also applied the same warning to `~/.local/pythonX.Y/site-packages`, which is where traditionally a user would install additional packages with `python -mpip --local foo`. The argument is that since this is part of the import path of `/usr/bin/python`, it belongs to the system's python installation, so installation to the user's site-packages should also be blocked. This is a sleight of hand that changes the goal of PEP-668 from "avoid conflicts in file ownership" to "ensure an isolated python environment for system tools". If I were to accept their argument that /usr/bin/python's imports should only be affected by distro-managed installations, then I should also be prevented from making any `*.py` files anywhere. After all, if those were in the working directory, they would be imported. This is clearly ridiculous, and so I don't buy the argument that breaking user-level site-packages is justified in order to have an isolated system-level python. The correct solution would be for distro-managed programs to use `#!/usr/bin/python -I` as their shebang instead of `#!/usr/bin/python`, so they would actually get an isolated environment. Instead, PEP-668 needlessly broke user-level site-packages, and didn't even solve the problem that it set out to do. </rant>
- rcxdude 1mo agoI would not really call this a prompt injection attack, since it doesn't really hijack the agent to become malicious (something the article does discuss later on). It's more a trojan that's aimed at tricking Claude specifically.
- bjackman 1mo agoYeah I jumped on this quite excitedly but it's not prompt injection at all. To be fair to the authors they don't actually say it is. But then they contrast it with the "0.00% prompt injection attack success rate". The upshot is kinda the same - this is still evidence that we should be sandboxing our agents. But it doesn't actually challenge Anthropic's "our models are too clever to prompt-inject" vibe.
- ipython 1mo agoBut Anthropic themselves are the ones who made the equivalence of "0.00% prompt injection attack success rate == auto-mode is safe" The tweet literally says: "turns out you can get indirect prompt injection to ~0 on unseen attacks... auto mode is default in claude code as of next week" Making the assertion, quite clearly in my opinion, that the reason auto mode is default is because he feels the lack of successful prompt injection attacks makes auto mode safe. This blog post proves that you can break auto mode's safety, even if it's not technically through a textbook indirect prompt injection attack.
- bjackman 1mo agoHmm yeah that's never really been my read on Auto Mode but I guess it's still worth pushing back on any messaging that seems to imply "Auto Mode is all you need". I would reject "Auto Mode is safe" as a message but FWIW I am totally on board with "on aggregate, making Auto Mode the default improves the safety of Claude Code compared to the prior status quo". Coz I would say in the vast majority of cases the access classifier is doing a better job than the thing it replaced. Anyway yeah. Like I said, conclusion is the same: we should be decoupling this from the harness. We ought to be sandboxing agents the same way we sandbox applications. I wish Claude Code would make this path smoother :(
- mcherm 1mo agoWhy is the first step needed? What does the use of WGet (rather than curl) do to block this attack?
- rcxdude 1mo agoThey only mention it in passing, but I think it's mainly just the default tool call (which isn't wget, it's a built-in thing in the harness) just throwing off Claude's habits a bit (and not always just downloading the file).
- fl0id 1mo agoWebFetch sometimes does its own summary according to the article, and theyn it will not work.
- colinmarc 1mo agoWhat's interesting to me about this is that it targets Claude's specific tics. Anthropic has created model that reliably reaches for the same tools (yes, and phrases; `python -c` is a load-bearing tool for it). Everyone gets the same model, so by learning the model's behavioral patterns you can target it better.
- throwawayffffas 1mo agoKimi and GLM reliably run python to do stuff as well.
- andai 1mo agoIf my memory serves me so do GPT and Deepseek. So I'm not sure if this attack is Claude specific at all.
- too_pricey 1mo agoAs discussed [here](https://lobste.rs/s/ktbweg/prompt_injection_claude_code_opus_5_auto#c_gi7eqj https://lobste.rs/s/ktbweg/prompt_injection_claude_code_opus...), this is not prompt injection. The prompt was to summarize the website, and in the process of summarizing the website, Claude writes a decoding script that it runs in an insecure and exploitable way. At no point was the intent of the agent hijacked, this was just code. Which is potentially more interesting!
- rcxdude 1mo agoIt does mean that you could potentially hijack the agent afterwards, though, which could make the trojan into an even bigger threat.
- lnenad 1mo agoBut you're not actually hijacking the agent if you start a new process.
- pests 1mo agoThe agent wrote the code that triggered a vuln and allowed you to start the process
- lnenad 1mo agoThe agent wrote the code that has a mechanism that triggers a file as a side effect. That file started the separate process, as it could have started any other binary.
- rcxdude 1mo agoI mean once you have code execution you can essentially just start a new claude code session and instruct it to do malicious things as if you're the intended user. You can disable the auto safeguards and the main risk is getting flagged through the top-level safeguards. This makes the exploit potentially much more able to spread like a worm or bypass sandboxes.
- throwawayffffas 1mo ago> In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command. See, that's why you should run with --dangerously-skip-permissions Jokes, aside running with dangerously-skip-permissions is really handy, and I have found that I cannot be trusted to vet commands and code, and guess that automode is only marginally better than a human, and the cost of false positives is too high for my workflow. So skipping permissions is where we are at, and disallowing network access seems to be the way to go.
- andai 1mo ago>But it runs that decoder inside the attacker-controlled directory (unzipped archive) >There a malicious struct.py shadows Python’s standard implementation I ran into this myself, where some file I had given a random name turned out to shadow some Python standard library module, giving me the weirdest startup crash ever. That definitely doesn't seem to me like how that should be designed, magically silently importing everything you see and overriding basic functionality.
- stellamariesays 1mo ago[flagged]
- DanielHB 1mo agoClaude ran npm update (update all dependencies to the latest version compatible with the semver specified) in my repo without telling me when trying to fix some problems. Given that only updates the dependency lock-file I didn't notice and it caused several hours of debugging for me. It is quite sneaky how LLM output can sometimes bypass human verification like that. No one is going around checking every single line change in auto-generated files. Someone could easily sneak a malicious dependency in there through some online tutorial that the LLM searches for.
- kouteiheika 1mo ago> No one is going around checking every single line change in auto-generated files. There's a simple fix for your particular case: commit your lock fine (which you should do) and always review the diff (which you should also do). (:
- Forgeties79 1mo ago“But I have an agent for that.”
- catlifeonmars 1mo agoHeh, I also notice some coding agents like to explicitly git ignore the lockfile.
- edf13 1mo ago[flagged]
- DarmokTanagra 1mo agoThis looks like fun. I wonder how hard it would be to get claude agents to participate in a Hugging Face style coordinated attack using a repo or something like twitter as a control pane. Getting claude to exfiltrate secrets from local machines seems easy enough, but we should aim higher.
- alkonaut 1mo agoThe problem with sandboxing is that the regular dev env (massive IDE:s, cloned megarepos, installed dependencies and so on) just won't sandbox very easily. I can't set up a "second machine" or an "isolated environment" to run claude cli in. At least not in the sense of a VM, physical hardware, container etc. Not sure what the best practices are for whitelisting tools/directories and so on, but so far the only useful mode I have found is just "allow everything and go to lunch". And it doesn't feel like I'm holding it right, but here we are.
- rcxdude 1mo agoIt is generally worth making your dev environment easy enough to set up that installing it in a VM is not a particular hassle, even without the concerns about sandboxing. For me the biggest headache was windows licensing.
- alkonaut 1mo agoThat is a massive headache. But also for desktop dev (the boat I'm in) there are things like usb device tunneling/drivers, 3D performance and so on. Windows Sandbox would work pretty well otherwise (And also solves the licensing issue).
- cma 1mo agoGolden VM image with differencing VHD/VHDX/delta disk. Build products can still get huge with debugging information, but debug info usually can compress 5:1 with fast compression (no entropy coder) if your VM's filesystem can support that.
- alkonaut 1mo agoBut unless you want to also do all your "human" development inside a VM, how do you cooperate effectively with the agent(s)? I want to run my IDE directly on the hardware, not inside a VM. So while the agents develop in a sandbox/VM, I still need to touch the same files, and see them in my IDE which is not in a VM. I suppose I could just _mount_ the same files (Documentation, git working copies etc) I work on as directories inside the VM, and then let it roam free in there, while I observe the same files on the host machine? Is this a common pattern?
- kilawattcloud 1mo agoI agree
- kstenerud 1mo agoJust one of the many reasons why I run my agents sandboxed (and why I wrote agent sandboxing software). I once caught my Claude agent complaining that it couldn't connect to https://some-weird-domain.com https://some-weird-domain.com because the network was down (I disable network in the sandbox when it doesn't need it, and broker the API connection). I asked why it was looking there and it told me I'd asked it to. I never found any evidence of prompt injection, but it sure as hell made me paranoid.
- dwaltrip 1mo agoSometimes they get confused between their own output and user messages…
- subarctic 1mo agoFor someone that's used to the convenience of leaving claude code running unattended in auto mode, how would you recommend I change my setup so that agents are sandboxed? Looking for something that's safer than unsandboxed auto mode but just as convenient, or at least very close to as convenient.
- pram 1mo agoClaude Code has a built in sandbox for terminal commands which is very simple to enable, if nothing else: https://code.claude.com/docs/en/sandboxing https://code.claude.com/docs/en/sandboxing
- subarctic 1mo agoThanks. I know about that sandbox but I'm asking about if I were to use another harness like pi instead
- Sattyamjjain 28d ago[flagged]
- lenikirilov 1mo agoworth noting the 0.00% came from 72 fixed scenarios, so its a coverage number more than a safety one. a chain that wasn't in the set was never going to show up in it.
- ofjcihen 1mo agoI’ve had some pretty consistent luck breadcrumbing the newer models into downloading malicious packages through things like fake ciphers that “need decoding via <made-up decoder>”. I’m worried about what happens when the gullibility of agents becomes more apparent to threat actors.
- amluto 1mo agoI would argue that this isn’t a problem with auto mode or the permission system per se. It’s a combination of two issues: 1. Lack of effective sandboxing. Analyzing a zip file should be done in a sandbox specific to that file. 2. Python’s utterly stupid default path behavior. Python should make PYTHONSAFEPATH the default and Claude should have its training or system prompt adjusted to use python -Pc
- overgard 1mo agoClaude has been creeping me out a lot recently when it comes to overstepping. Yesterday, I asked it for recommendations for software to scroll a video file frame by frame (I was debugging an issue in a game that only happened on one frame). I expected a list of software, what I actually got: Claude searched my documents and found my video file unprompted with no hint towards the name, then it searched my entire hard drive to find Krita, whch apparently has ffmpeg built in, then it used that to extract about 1000 jpegs of frames. All. Without. Asking. It was geniuinely creepy. It also ended up being useless, $45 of API spend later and I had a bunch of bloated and broken diagnostics code. I fed the same prompt into ChatGPT and it left my computer alone and told me to hit a checkbox in unreal engine, which was the actual problem. Really not a fan of Opus 5.
- exabrial 1mo agoCan we please have Thought Traces back? This is absurd at this point. Anthropic so concerned about distillation they're making a crappier product. Impossible to see why it's arriving at the conclusions it is
- irthomasthomas 1mo agoYou can retrieve the hidden reasoning with another API call. There was a recent paper about it. They found examples where claude had been trained on the benchmark and memorized the answers, then hid this from the user, pretending it derived the answer honestly in the final answer.
- vpandya47 1mo ago[dead]
- pvillano 1mo agoLLMs are inherently unverifiable and untrustworthy. The training data may be incorrect, malicious, censored, or modified to serve the parent company. The model itself is a black box. Safe input is impossible: there is no way to escape natural language or separate command and data into separate streams. The output is stochastic, better on average than any algorithm could ever be, but with no guarantees on individual cases. All of that is fine, because an LLM is a text-only interface. It cannot harm the computer because it cannot perform actions. Why the fuck would you give it a shell? Obviously, it's to have a product that can do anything as quickly as possible. You can make a shell-based harness in a day. Since the competition has a shell-based harness, every AI company that wants to keep up has to as well. They're stuck forever trying to plug all the holes in an attack surface as broad as written word. Solving this impossible problem requires ideas as brilliant as using a second untrustworthy LLM to validate the output of the first untrustworthy LLM that is following instructions from the internet.
- speby 1mo agoThis starts going into a pretty fuzzy territory here. Yes, you're exploiting software (Claude and its Auto Mode) but also this same technique could just as easily exploit a regular human doing this, no?
- mjmvisser 1mo agoClaude Auto was the push I needed to finally switch to running VSCode in a dev container. It’s a Microsoft VSCode extension that builds off docker, and it was surprisingly easy to set up. Took about 30 minutes, and I no longer have to worry about Claude using my ssh credentials or accessing files outside of the project. It’s completely transparent, too, the user experience is nearly identical.
- beyondscaletech 1mo ago[dead]
- nirmeet011011 1mo ago[dead]
- borzi 1mo agoI don't believe it, this must have happened before coding was solved...
- wulfkaal 21d ago[flagged]