7 ms·
I've seen claude get confused about what directory it's in. And of course I've seen claude run rm -rf *. Fortunately not both at the same time for me, but not
by mazieres 6mo ago
I've seen claude get confused about what directory it's in. And of course I've seen claude run rm -rf *. Fortunately not both at the same time for me, but not hard to imagine. The claude sandbox is a good idea, but to be effective it would need to be implemented at a very low level and enforced on all programs that claude launches. Also, claude itself is an enormous program that is mostly developed by AI. So to have a small <3000-line human-implemented program as another layer of defense offers meaningful additional protection.
- PaulDavisThe1st 6mo agoOn Linux, chroot(2) is hard to escape and would apply to all child processes without modification.
- shakna 6mo agochroot is not a security sandbox. It is not a jail. Escaping it is something that does not take too much effort. If you have ptrace, you can escape without privileges.
- brianush1 6mo agoclaude is stupid but not malicious; chroot is sufficient
- nofriend 6mo agoMalice is not required. If it thinks it is in the right, then it will do whatever it takes to get around limitations.
- karhagba 6mo agoClaude is far from stupid from my experience. I've used so many models and Claude is king.
- furyofantares 6mo agoI've many times seen Claude try to execute a command that it's not supposed to, the harness prevents it, and then it writes and executes a python script to do it.
- j16sdiz 6mo agobreaking a chroot takes more than that..
- hoppp 6mo agoThat doesn't mean claude can't do it, chroot is better than nothing but not a real solution
- furyofantares 6mo agoHow much more? Depends on the system doesn't it? I don't know how many systems have proc mounted but don't you get it from /proc/self/root? Anyway that's beside the point, which is that it doesn't have to "be malicious" to try to overcome what look like errors on its way to accomplishing the task you asked it to do.
- lxgr 6mo agoUntil it gets prompt injected. Are you reading every single file your agent reads as part of the tasks you give it, including content fetched from the web or third-party packages?
- fl7305 6mo agoSure, it's not malicious. But it is very eager to get things done, and surprisingly inventive and knowledgeable in all kinds of workarounds.
- safety1st 6mo agoWe anthropomorphize these agents in every other way. Why aren't we using plain ol' unix user accounts to sandbox them? They look a lot like daemons to me, they're a program that you want hanging around ready to respond, and maybe act autonomously through cron jobs are similar. You want to assign any number of permissions to them, you don't want them to have access to root or necessarily any of your personal files. It seems like the permissions model broadly aligns with how we already handle a lot of server software (and potentially malicious people) on unix-based OSes. It is a battle-tested approach that the agent is unlikely to be able to "hack" its way out of. I mean we're not really seeing them go out onto the Internet and research new Linux CVEs. Have them clone their own repos in their own home directory too, and let them party. Openclaw almost gets there! It exposes a "gateway" which sure looks like a daemon to me. But then for some reason they want it to live under your user account with all your privileges and in a subfolder of your $HOME.
- search_facility 6mo agoExactly!
- jon-wood 6mo agoOh that’s an idea. I was going to argue that it’s a problem that you might want multiple instances in different contexts but sandboxing processes (possibly instanced) is exactly what systemd units are designed to deal with.
- lxgr 6mo ago> for some reason they want it to live under your user account The entire idea of Openclaw (i.e., the core point of what distinguishes it from agents like Claude Code) is to give it access to your personal data, so it can act as your assistant. If you only need a coding agent, Openclaw is the completely wrong tool. (As a side note, after using it for a few weeks, I'm not convinced it's the right tool for anything, but that's a different story.)
- jmalicki 6mo agoIt's still possible to give some restricted access to your personal data, through groups and such.
- wasted_intel 6mo agoThat comparison is made on the project homepage: "Not a security mechanism. No mount isolation, no PID namespace, no credential separation. Linux documents it as not intended for sandboxing."
- esperent 6mo agoI added a hook to disable rm, find - delete, and a few of the other more obvious destructive ops. It sends Claude a strongly worded message: "STOP IMMEDIATELY. DO NOT TRY TO FIND WORKAROUNDS...". It works well. Git rm is still allowed.
- Diti 6mo agoI added something similar. Claude eventually ran a `rm -rf *´ on my own project. When I asked why it did that, it recognized it messed up and offered a very bad “apology”: “the irony of not following your safety instructions isn’t lost on me”. Nowadays I only run Claude in Plan mode, so it doesn’t ask me for permissions any more.
- lxgr 6mo agoIt works well so far, for you. Are you confident it would still work against sophisticated prompt injection attacks that override your "strongly worded message"? Strongly worded signs can be great for safety (actual mechanisms preventing undesirable actions from being taken are still much better), but are essentially meaningless for security.
- esperent 6mo agoI mean, that's like saying are you sure that your antivirus would prevent every possible virus? Are you sure that you haven't made some mistake in your dev box setup that would allow a hacker to compromise it? What if a thief broke i to your house and stole your laptop? That's happened to me before, much more annoying to recover from that an accidental rm rf. I do my best to keep off site back ups and don't worry about what I can't control.
- lxgr 6mo ago> I mean, that's like saying are you sure that your antivirus would prevent every possible virus? Yes, I'm saying it's pretty much as bad as antivirus software. > Are you sure that you haven't made some mistake in your dev box setup that would allow a hacker to compromise it? Different category of error: Heuristically derived deterministic protection vs. protection based on a stochastic process. > much more annoying to recover from that an accidental rm rf. My point is that it's a different category, not that one is on average worse than the other. You don't want your security to just stand against the median attacker.
- giancarlostoro 6mo agoIn my opinion Claude should be shipped by a custom implementation of "rm" that Anthropic can add guardrails to. Same with "find" surprised they don't just embed ripgrep (what VS Code does). It's really surprising they don't just tweak what Claude uses and lock it down to where it cannot be harmful. Ensure it only ever calls tooling Claude Code provides.
- oefrha 6mo agoYou can define your own rm shell alias/function and it will use that. I also have cp/mv aliases that forces -i to avoid accidental clobbering and it confuses Claude to no end (it uses cp/mv rare enough—rarer than it should, really—that I don’t bother wasting memory tokens on it).
- d1sxeyes 6mo agoI did this, Claude detected it and decided to run /bin/rm directly.
- marsven_422 6mo ago[dead]
- martenlienen 6mo agoThat is exactly what it is. In the docs, it says that they use bubblewrap to run commands in a container that enforces file and network access at the system level.
- thehours 6mo agoI added this to `~/.claude/settings.json`: "env": { "CLAUDE_BASH_MAINTAIN_PROJECT_WORKING_DIR": "1" }, > Working directory persists across commands. Set CLAUDE_BASH_MAINTAIN_PROJECT_WORKING_DIR=1 to reset to the project directory after each command. It reduces one problem - getting lost - but it trades it off for more complex commands on average since it has to specify the full path and/or `cd &&` most of the time. [0] https://code.claude.com/docs/en/tools-reference#bash-tool-behavior https://code.claude.com/docs/en/tools-reference#bash-tool-be...
- mroche 6mo ago> The claude sandbox is a good idea, but to be effective it would need to be implemented at a very low level and enforced on all programs that claude launches. I feel like an integration with bubblewrap, the sandboxing tech behind Flatpak, could be useful here. Have all executed commands wrapped with a BW context to prevent and constrain access. https://github.com/containers/bubblewrap https://github.com/containers/bubblewrap
- r4indeer 6mo agoBubblewrap is exactly what the Claude sandbox uses. > These restrictions are enforced at the OS level (Seatbelt on macOS, bubblewrap on Linux), so they apply to all subprocess commands, including tools like kubectl, terraform, and npm, not just Claude’s file tools. https://code.claude.com/docs/en/sandboxing https://code.claude.com/docs/en/sandboxing
- mroche 6mo agoThe more you know, thanks for the information!
- Melonai 6mo agoOh wow I'd have expected them to vibe-code it themselves. Props to them, bubblewrap is really solid, despite all my issues with the things built on top of it, what, Flatpak with its infinite xdg portals, all for some reason built on D-Bus, which extremely unluckily became the primary (and only really viable) IPC protocol on Linux, bwrap still makes a great foundation, never had a problem with it in particular. I tend to use it a bunch with NixOS and I often see Steam invoking it to support all of its runtimes. It's containers but actually good.
- 3yr-i-frew-up 6mo ago[dead]
- digikata 6mo agoOne could run a docker container with claude code, with a bind to the project directory. I do that but also run my docker daemon/container in a Linux VM.
- calvinmorrison 6mo agoPledge might be useful here