11 ms·
Go hard on agents, not on your filesystem
- mazieres 6mo agoWhat would it take for people to stop recklessly running unconstrained AI agents on machines they actually care about? A Stanford researcher thinks the answer is a new lightweight Linux container system that you don't have to configure or think about.
- mememememememo 6mo agoYes. It is like walking arounf your house with a flamethrower, but you added fire retardant. Just take the flamethower to a shed you don't mind losing. Which is some kind of cloud workspace most likely. Maybe an old laptop. Still if you yolo online access and give it cred or access to tools that are authenticated there can still be dragons.
- cindyllm 6mo ago[dead]
- mazieres 6mo agoThe problem is that in practice, many people don't take the flamethrower to the shed. I recently had a conversation with someone who was arguing that you don't really need jai because docker works so well. But then it turned out this person regularly runs claude code in yolo mode without a container! It's like people think that because containers and VMs exist, they are probably going to be using them when a problem happens. But then you are working in your own home directory, you get some compiler error or something that looks like a pain to decipher, and the urge just to fire up claude or codex right then and there to get a quick answer is overwhelming. Empirically, very few people fire up the container at that point, whereas "jai claude" or "jai -D claude" is simple enough to type, and basically works as well as plain claude so you don't have to think about it.
- fouc 6mo agoexcept the big AI companies are pushing stuff designed for people to run on their personal computers, like Claude Cowork.
- vardalab 6mo agounconstrained AI agents are what makes it so useful though. I have been using claude for almost a year now and the biggest unlock was to stop being a worrywart early on and just literally giving it ssh keys and telling it to fix something. ofc I have backups and do run it in VM but in that VM it helps me manage by infra and i have a decent size homelab that would be no fun but a chore without this assistant.
- kristofferR 6mo agoAgree, but SSH agents like 1Passwords are nice for that. You simply tell it to install that Docker image on your NAS like normal, but when it needs to login to SSH it prompts for fingerprint. The agent never gets access to your SSH key.
- sersi 6mo agoI run my AI agent unconstrained in a VM without access to my local network so it can futz with the system however it wants (so far, I've had to rebuild the VM twice from Claude borking it). That works great for software development. For devops work, etc (like your use case), I much prefer talking to it and letting it guide me into fixing the issue. Mostly because after that I really understand what the issue was and can fix it myself in the future.
- bigstrat2003 6mo ago> unconstrained AI agents are what makes it so useful though Not remotely worth it.
- hrmtst93837 6mo ago[flagged]
- throwingSE 6mo agojai is doing the right thing for its threat model. The credential layer is a different surface though ... an agent with a broad API token can call initiate_payment or update_vendor_bank on a remote production system and the filesystem sandbox can't help. Applying the same principle as jai for remote boundaries, we can scope API authority to the task
- jillesvangurp 6mo agoThere always has been this tension between protecting resources and allowing users to access those resources in security. With many systems you have admin/root users and regular users. Some things require root access. Most interesting things (from a security point of view) live in the user directory. Because that's where users spend all their time. It's where you'll find credentials, files with interesting stuff inside, etc. All the stuff that needs protecting. The whole point of using a computer is being able to use it. For programmers, that means building software. Which until recently meant having a lot of user land tools available ready to be used by the programmer. Now with agents programming on their behalf, they need full access to all that too in order to do the very valuable and useful things they do. Because they end up needing to do the exact same things you'd do manually. The current security modes in agents are binary. Super anal about absolutely everything; or off. It's a false choice. It's technically your choice to make and waive their liability (which is why they need you to opt in); but the software is frustrating to use unless you make that choice. So, lots of people make that choice. I'm guilty as well. I could approve every ansible and ssh command manually (yes really). But a typical session where codex follows my guardrails to manage one of my environments using ansible scripts it maintains just involves a whole lot such commands. I feel dirty doing it. But it works so well that doing all that stuff manually is not something I want to go back to. It's of course insecure as hell and I urgently need something better than yolo mode for this. One of the reasons I like codex is that (so far) it's pretty diligent about instruction following and guard rails. It's what makes me feel slightly more relaxed than I perhaps should be. It could be doing a lot of damage. It just doesn't seem to do that.
- BoppreH 6mo agoExcellent project, unfortunate title. I almost didn't click on it. I like the tradeoff offered: full access to the current directory, read-only access to the rest, copy-on-write for the home directory. With stricter modes to (presumably) protect against data exfiltration too. It really feels like it should be the default for agent systems.
- fouc 6mo agoSince the site itself doesn't really have a title, I probably would've went with something like "jai - filesystem containment for AI agents"
- drtournier 6mo ago[flagged]
- mememememememo 6mo agoSo?
- triilman 6mo agoWhat would Jonathan Blow think about this.
- ghighi7878 6mo agoMy name is also jai
- messh 6mo agoHow is this different than say bubblewrap and others?
- girvo 6mo agohttps://jai.scs.stanford.edu/comparison.html#jai-vs-bubblewrap https://jai.scs.stanford.edu/comparison.html#jai-vs-bubblewr... > bubblewrap is more flexible and works without root. jai is more opinionated and requires far less ceremony for the common case. The 15-flag bwrap invocation that turns into a wrapper script is exactly the friction jai is designed to remove. Plus some other comparisons, check the page
- attentive 6mo agobubblewrap is in many modern distros standard packages. With all the supply chain issues these days onboarding new tools carries extra risks. So, question is if it's worth it.
- deleted 6mo ago[deleted]
- simonw 6mo agoSuggestion for the FAQ page: does this work on a Mac?
- adi_kurian 6mo agoClaude's stock unprompted / uninspired UI code creates carbon clone components. That "jai is not a promise of perfect safety" callout box is like the em dash of FE code. The contrast, or lack thereof, makes some of the text particularly invisible. I wonder if shitty looking websites and unambitious grammar will become how we prove we are human soon.
- NetOpWibby 6mo agoEverything old is new again
- deleted 6mo ago[deleted]
- AnotherGoodName 6mo agoAdd this to .claude/settings.json: { "sandbox": { "enabled": true, "filesystem": { "allowRead": ["."], "denyRead": ["~/"], "allowWrite": ["."], "denyWrite": ["/"] } } } You can change the read part if you're ok with it reading outside. This feature was only added 10 days ago fwiw but it's great and pretty much this.
- mycall 6mo agoI noticed codex has a sandbox, wondering if it has a comparable config section.
- tofflos 6mo agoCodex uses and ships with bubblewrap on Linux and will attempt to use the version installed on the path before falling back to the shipped version with a warning message. You should be able to configure the sandbox using https://developers.openai.com/codex/agent-approvals-security https://developers.openai.com/codex/agent-approvals-security if you are a person who prefers the convenience of codex being able to open the sandbox over an externally enforced sandbox like jai.
- harikb 6mo agoI think the point would be that - some random upcoming revision of claude-code could remove or simply change the config name just as silently as it was introduced. People might genuinely want some other software to do the sandboxing. Something other than the fox.
- cozzyd 6mo agoIs this a real sandbox or just a pretty please?
- gerdesj 6mo ago[flagged]
- cozzyd 6mo agoShould be named Jia More seriously, I'm not a heavy agent user, but I just create a user account for the agent with none of my own files or ssh keys or anything like that. Hopefully that's safe enough? I guess the risk is that it figures out a local privilege escalation exploit...
- timcobb 6mo agoDunno... with this setup it seems certain that the agent will discover a zero-day to escalate privilges and send your SSH keys to its handlers in N. Korea. P.S. Everything old is new again <3
- cozzyd 6mo agoYeah definitely a concern. Probably need a sandbox and separate user for defense in depth.
- mbreese 6mo agoThis still is running in an isolated container, right? Ignoring the confidentiality arguments posed here, I can’t help to think about snapshotting filesystems in this context. Wouldn’t something like ZFS be an obvious solution to an agent deleting or wildly changing files? That wouldn’t protect against all issue the authors are trying to address, but it seems like an easy safeguard against some of the problems people face with agents.
- gurachek 6mo agoThe examples in the article are all big scary wipes, But I think the more common damage is way smaller and harder to notice. I've been using claude code daily for months and the worst thing that happened wasnt a wipe(yet). It needed to save an svg file so it created a /public/blog/ folder. Which meant Apache started serving that real directory instead of routing /blog. My blog just 404'd and I spent like an hour debugging before I figured it out. Nothing got deleted and it's not a permission problem, the agent just put a file in a place that made sense to it. jai would help with the rm -rf cases for sure but this kind of thing is harder to catch because its not a permissions problem, the agent just doesn't know what a web server is.
- cozzyd 6mo agoShould definitely block .ssh reading too...
- justinde 6mo ago.claude/settings.json: { "sandbox": { "enabled": true, "filesystem": { "allowRead": ["."], "denyRead": ["~/"], "allowWrite": ["."] } } } Use it! :) https://code.claude.com/docs/en/sandboxing https://code.claude.com/docs/en/sandboxing
- charcircuit 6mo agoI want agents to modify the file system. I want them to be able to manage my computer if it thinks it's a good idea. If a build fails due to running out of disk space I want it to be able to find appropriate stuff to delete to free up space.
- gonzalohm 6mo agoNot sure I understand the problem. Are people just letting AI do anything? I use Claude Code and it asks for permission to run commands, edit files, etc. No need for sandbox
- mazieres 6mo agoYes, people very much are, and that's exactly the problem! People run `claude --dangerously-skip-permissions` and `codex --yolo` all the time. And I think one of the appeals of opencode (besides cross-model, which is huge) is that the permissions are looser by default. These options are presumably intended for VM or container environments, but people are running them outside. And of course it works fine the first 100 times people do it, which drives them to take bigger and bigger risks.
- kristofferR 6mo agoAlso recommended: https://github.com/kenryu42/claude-code-safety-net https://github.com/kenryu42/claude-code-safety-net
- Jach 6mo agoI've done some experimenting with running a local model with ollama and claude code connecting to it and having both in a firejail: https://firejail.wordpress.com/ https://firejail.wordpress.com/ What they get access to is very limited, and mostly whitelisted.
- e1g 6mo agoFor jailing local agents on a Mac, I made Agent Safehouse - it works for any agent and has many sane default for developers https://agent-safehouse.dev https://agent-safehouse.dev
- ray_v 6mo agoI'm wondering if the obvious (and stated) fact that the site was vibe-coded - detracts from the fact that this tool was hand written. > jai itself was hand implemented by a Stanford computer science professor with decades of C++ and Unix/linux experience. (https://jai.scs.stanford.edu/faq.html#was-jai-written-by-an-ai-coding-agent https://jai.scs.stanford.edu/faq.html#was-jai-written-by-an-...)
- Quarrel 6mo agoTo be less abstract, it was written by David Mazieres, who was been writing software and papers about user level filesystems since at least 2000. He now runs the Stanford Secure Computer Systems group. David has done some great work and some funny work. Sometimes both.
- mazieres 6mo agoHuman author here. The fact that I don't know web design shouldn't detract from my expertise in operating systems. I wrote the software and the man page, and those are what really matter for security. The web site is... let's say not in a million years what I would have imagined for a little CLI sandboxing tool. I literally laughed out loud when claude pooped it out, but decided to keep, in part ironically but also since I don't know how to design a landing page myself. I should say that I edited content on the docs part of the web site to remove any inaccuracies, so the content should be valid.
- Nifty3929 6mo agoIndeed! Kinda reminds me of this: https://m.xkcd.com/932/ https://m.xkcd.com/932/ I'm not a web UI guy either, and I am so, so happy to let an AI create a nice looking one for me. I did so just today, and man it was fast and good. I'll check it for accuracy someday...
- lifis 6mo agoIt seems that the LLM has not only designed the site, but also written the text on at least the frontpage, which is a pretty bad signal. You need to rewrite all the text and Telde it with text YOU would actually write, since I doubt you would write in that style.
- rsyring 6mo agoI've been reviewing Agent sandboxing solutions recently and it occurred to me there is a gaping vector for persistent exploits for tools that let the agent write to the project directory. Like this one does. I had originally thought this would ok as we could review everything in the git diff. But, it later occurred to me that there are all kinds of files that the agent could write to that I'd end up executing, as the developer, outside the sandbox. Every .pyc file for instance, files in .venv , .git hook files. ChatGPT[1] confirms the underlying exploit vectors and also that there isn't much discussion of them in the context of agent sandboxing tools. My conclusion from that is the only truly safe sandboxing technique would be one that transfers files from the sandbox to the dev's machine through some kind of git patch or similar. I.e. the file can only transfer if it's in version control and, therefore presumably, has been reviewed by the dev before transfer outside the sandbox. I'd really like to see people talking more about this. The solution isn't that hard, keep CWD as an overlay and transfer in-container modified files through a proxy of some kind that filters out any file not in git and maybe some that are but are known to be potentially dangerous (bin files). Obviously, there would need to be some kind of configuration option here. 1: https://chatgpt.com/share/69c3ec10-0e40-832a-b905-31736d8a3438 https://chatgpt.com/share/69c3ec10-0e40-832a-b905-31736d8a34...
- mazieres 6mo agoIt's a good point. Maybe I should add an option to make certain directories read-only even under the current working directory, so that you can make .git/ read-only without moving it out of the project directory. You can already make CWD an overlay with "jai -D". The tricky part is how to merge the changes back into your main working directory.
- avazhi 6mo ago[flagged]
- faangguyindia 6mo agoi just use seatbelt (mac native) in my custom coding agent: supercode
- stavros 6mo agoI'd really like to try this, but building it is impossible. C++ is such a pain to build with the "`make`; hunt for the dependency that failed; `apt-get install whatever-dev`; goto make" loop... Please release binaries if you're making a utility :(
- jbverschoor 6mo agohttps://github.com/jrz/container-shell https://github.com/jrz/container-shell It does something very simple, and it’s a POSIX shell script. Works on Linux and macOS. Uses docker to sandbox using bind mount
- stavros 6mo agoYeah but it doesn't COW anything else, and Docker is a bit heavy for this.
- mazieres 6mo agoWhat distro are you using? The only two dependencies are libacl and libmount. I'm trying to figure out which distros don't include these by default, and if the libraries are really missing, or if it's just the pkgconf ".pc" files. In the former case I should document the dependencies. In the latter case I should maybe switch from PKG_CHECK_MODULES to old-fashioned autoconf.
- stavros 6mo agoI'm using Ubuntu, I gave up when it failed on something about "print".
- jbverschoor 6mo agoInteresting take on the same problem I created https://github.com/jrz/container-shell https://github.com/jrz/container-shell which basically launches a persistent interactive shell using docker, chrooted to the CWD CWD is bind mounted so the rest is simply not visible and you can still install anything you want.
- waterfisher 6mo agoThere's nothing wrong with an AI-designed website, but I wish when describing their own projects that HN contributors wrote their own copy. As HN posters are wont to say, writing is thinking...
- rdevsrex 6mo agoThis won't cause any confusion with the jai language :)
- Waterluvian 6mo agoAre mass file deletions as result of some plausible “I see why it would have done that” or will it just completely randomly execute commands that really have nothing to do with the immediate goal?
- deleted 6mo ago[deleted]
- puttycat 6mo agoI am still amazed that people so easily accepted installing these agents on private machines. We've been securing our systems in all ways possible for decades and then one day just said: oh hello unpredictable, unreliable, Turing-complete software that can exfiltrate and corrupt data in infinite unknown ways -- here's the keys, go wild.
- fc417fc802 6mo agoPeople were also dismissing concerns about build tooling automatically pulling in an entire swarm of dependencies and now here we are in the middle of a repetitive string of high profile developer supply chain compromises. Short term thinking seems to dominate even groups of people that are objectively smarter and better educated than average.
- culopatin 6mo agoIf anything I feel more in control of these agents than the millions of LOC npm or pip pull in to just show me a hello world
- Sindisil 6mo agoThe load bearing word being "feel".
- tokioyoyo 6mo ago> “high profile developer supply chain compromises” And nothing big has happened despite all the risks and problems that came up with it. People keep chasing speed and convenience, because most things don’t even last long enough to ever see a problem.
- fc417fc802 6mo agoI've yet to be saved by an airbag or seatbelt. Is that justification to stop using them? How near a miss must we have (and how many) before you would feel that certain practices surrounding dependencies are inadvisable? A number of these supply chain compromises had incredibly high stakes and were seemingly only noticed before paying off by lucky coincidence.
- andai 6mo agoThis looks great and seems very well thought out. It looks both more convenient and slightly more secure than my solution, which is that I just give them a separate user. Agents can nuke the "agent" homedir but cannot read or write mine. I did put my own user in the agent group, so that I can read and write the agent homedir. It's a little fiddly though (sometimes the wrong permissions get set, so I have a script that fixes it), and keeping track of which user a terminal is running as is a bit annoying and error prone. --- But the best solution I found is "just give it a laptop." Completely forget OS and software solutions, and just get a separate machine! That's more convenient than switching users, and also "physically on another machine" is hard to beat in terms of security :) It's analogous to the mac mini thing, except that old ThinkPads are pretty cheap. (I got this one for $50!)
- lll-o-lll 6mo agoWhere this falls down is that for the agents to interact with anything external, you have to give them keys. Without a proxy handling real keys between your agent and external services, those keys are at risk of compromise. Also. Agents are very good at hacking “security penetration testing”, so “separate user” would not give me enough confidence against malicious context.
- sanitycheck 6mo agoSo don't let them interact with anything external. You can push and pull to their git project folders over the local filesystem or network, they don't even need access to a remote.
- lll-o-lll 6mo agoUnless you are talking about running a local model, that’s not possible.
- sanitycheck 6mo agoObviously if you're running Claude Code you need a token for that and an internet connection, that's kind of a given. What I'm talking about is permission (OS level, not a leaky sandbox) to access the user's files, environment variables, project credentials for git remotes, signing keys, etc etc.
- samchon 6mo agoJust allowing Yolo, and sometimes do rolling back
- KennyBlanken 6mo agoThis is not some magical new problem. Back your shit up. You have no excuse for "it deleted 15 years of photos, gone, forever."
- sersi 6mo agoAnd what about, it exfiltrated my AWS keys (or insert random valuable thing that sits in .config of your home directory)? Backing up is not going to help you in that case.
- yalogin 6mo agoWhat if Claude needs me to install some software and hoses my distro. Jai cannot protect there as I am running the script myself
- schaefer 6mo agoUgh. The name jai is very taken[1]... names matter. [1]: https://en.wikipedia.org/wiki/Jai_(programming_language) https://en.wikipedia.org/wiki/Jai_(programming_language)
- vscode-rest 6mo agoSlightly taken, at best.
- diego_sandoval 6mo agoJonathan Blow has said that "Jai" is just a placeholder name or something.
- schaefer 6mo agoI hadn’t heard that. Thanks
- john_strinlai 6mo agoa closed beta of an obscure programming language where the wikipedia page is nominated for deletion because it is a "Non-notable programming language that is not publicly available." is considered "very taken"?
- qq66 6mo agoThat's an unreleased product in closed beta. Might not any name conflict with some unreleased product in closed beta?
- albert_e 6mo agoCan we have a hardware level implementation of git (the idea of files/data having history preserved. Not necessarily all bells and whistles.) ...in a future where storage is cheap.
- samlinnfer 6mo agoNow we just need one for every python package.
- gck1 6mo agoIt's full VM or nothing. I want AI to have full and unrestricted access to the OS. I don't want to babysit it and approve every command. Everything that is on that VM is a fair game and the VM image is backed up regularly from outside. This is the only way.
- griffindor 6mo agoI use Nix shells to give it the tools it wants. If it wants to do system-level tests, then I make sure my project has Qemu-based tests.
- adi_kurian 6mo agoI have a pretty insane thing where I patched the screen sharing binary and hand rolled a dummy MDN so I can have multiple profiles logged in at once on my Mac Studio. Then have screen share of diff profiles in diff "windows". Was for some ML data gathering / CV training. It's pretty neat, screen sharing app is extremely high quality these days, I can barely notice a diff unless watching video. Almost feels like Firefox containers at OS level. Have thought that could be a pretty efficient way to have restricted unrestricted convenient AI access. Maybe I'll get around to that one day.
- gck1 6mo ago> I have a pretty insane thing where I patched the screen sharing binary and hand rolled a dummy MDN so I can have multiple profiles logged in at once on my Mac Studio I have a Studio collecting dust that I've been eyeing every time my VM crashed because of Apple's paravirtualized GPU proxy not being able to keep up with things I run in it. This sounds exactly like what I wanted to do on my Studio and didn't know where to pull the thread from. Do you have this method shared openly anywhere?
- adi_kurian 6mo agoNah but I'd be happy to share it with you over DM! (If they have DMs on here?)
- gpm 6mo agoThis is a cool solution... I have a simpler one, though likely inferior for many purposes.. Run <ai tool of your choice> under its own user account via ssh. Bind mount project directories into its home directory when you want it to be able to read them. Mount command looks like sudo mkdir /home/<ai-user>/<dir-name> sudo mount --bind <dir to mount> --map-groups $(id -g <user>):$(id -g <ai-user>):1 --map-users $(id -u <user>):$(id -u <ai-user>):1 /home/<ai-user>/<dir-name> I particularly use this with vscode's ssh remotes.
- athrowaway3z 6mo agoI've been using a dedicated user account for 6 months now, and it does everything. What makes it great is the only axis of configuration is managing "what's hoisted into its accessible directories". Its awe-inspiring the levels of complexity people will re-invent/bolt-on to achieve comparable (if not worse) results.
- kevinbaiv 6mo ago[flagged]
- sanskritical 6mo agoHow long until agents begin routinely abusing local privilege escalation bugs to break out of containers? I bet if you tell them explicitly not to do so it increases the likelihood that they do.
- neilwilson 6mo agoIt's always struck me that agents should be operated via `systemd-run` as a transient scope unit with the necessary security properties set So couldn't this be done with an appropriate shell alias - at least under linux.
- _shadi 6mo agoI had the same idea and created this quickly in an evening: https://github.com/Shadi/isolate https://github.com/Shadi/isolate
- orthogonalinfo 6mo ago[flagged]
- ta-run 6mo agoIdk, just feels so counter sometimes to build and refine these (seemingly non-deterministic) tools to build deterministic workflows & get the most productivity out of them.
- 0xbadcafebee 6mo agoIf it has a big splash page with no technical information, it's trying to trick you into using it. That doesn't mean it isn't useful, but it does mean it's disingenuous. This particular solution is very bad. To start off with, it's basically offering you security, right? Look, bars in front of an evil AI! An AI jail! That's secure, right? Yet the very first mode it offers you is insecure. The "casual" mode allows read access to your whole home directory. That is enough to grant most attackers access to your entire digital life. Most people today use webmail. And most people today allow things like cookies to be stored unencrypted on disk. This means an attacker can read a cookie off your disk, and get into your mail. Once you have mail, you have everything, because virtually every account's password reset works through mail. And this solution doesn't stop AI exfiltration of sensitive data, like those cookies, out the internet. Or malware being downloaded into copy-on-write storage space, to open a reverse shell and manipulate your existing browser sessions. But they don't mention that on the fancy splash page of the security tool. The truth is that you actually need a sophisticated, complex-as-hell system to protect from AI attacks. There is no casual way to AI security. People need to know that, and splashy pages like this that give the appearance of security don't help the situation. Sure, it has disclaimers occasionally about it not being perfect security, read the security model here, etc. But the only people reading that are security experts, and they don't need a splash page! Stanford: please change this page to be less misleading. If you must continue this project with its obviously insecure modes, you need to clearly emphasize how insecure it is by default. (I don't think it even qualifies as security software)
- yobert 6mo agoIt is a bit better than you're saying. When you fire it up, you can see that it does have a list of common credential areas that it hides from the jail. It seems to hide: .aws .azure .bash_history .config .docker .git-credentials .gnupg .jai .local .mozilla .netrc .password-store .ssh .zsh_history It's a humorous attempt in a sense, but better than nothing for sure!
- deleted 6mo ago[deleted]
- techpulselab 6mo ago[flagged]
- hikaru_ai 6mo ago[dead]
- lemontheme 6mo agoAnd for the macos users, I can’t recommend nono enough. (Paying it forward, since it was here on HN that I learned about it.) Good DX, straightforward permissions system, starts up instantly. Just remember to disable CC’s auto-updater if that’s what you’re using. My sandbox ranking: nono > lima > containers.
- pbowyer 6mo agoThis nono? https://github.com/always-further/nono https://github.com/always-further/nono > Just remember to disable CC’s auto-updater if that’s what you’re using. Why?
- lemontheme 6mo agoMight be something specific to my and my colleagues' systems, but it breaks the TUI. It needs git authentication, which fails, and the TUI stops accepting input reliably
- faeyanpiraat 6mo agoI've just switched to lima, and cant find anything about "nono" can you post a link?
- lemontheme 6mo agoI really like lima too. It's my go-to recommendation for light VMs. But I do consider it slightly less convenient. A good example of why is project-local .venv/ directories, which are the default with uv. With Lima, what happens is that macOS package builds get mounted into a Linux system, with potential incompatibility issues. Run uv sync inside the VM and now things are invalid on the macOS side. I wasn't able to find a way to mount the CWD except for certain subdirectories. Another example is network filtering. Lima (understandably) doesn't offer anything here. You can set up a firewall inside the VM, but there's no guarantee your agent won't find a way to touch those rules. You can set it up outside the VM, but then you're also proxying through a MITM. So, for the use case of running Claude Code in --dangerously-skip-permissions mode, Lima is more hassle than Nono
- ozim 6mo agoI have seen it just 5 mins ago Claude misspelled directory path - for me it was creating a new folder but I can image if I didn’t stop it it could start removing stuff just because he thinks he needs to start from scratch or something.
- bob1029 6mo agoI've been running GPT5.x fully unconstrained with effective local admin shell for over $500 worth of API tokens. Not once has it done something I'd consider "naughty". It has left my project in a complete mess, but never my entire computer. git reset --hard && git clean -fd That's all it takes. I think this is turning into a good example of security theatrics. If the agent was actually as nefarious as the marketing here suggests, the solution proposed is not adequate. No solution is. Not even a separate physical computer. We need to be honest about the size of this problem. Alternatively, maybe Claude is unusually violent to the local file system? I've not used it at all, so perhaps I am missing something here.
- mazieres 6mo agoAn AI agent that works 99.9% of the time is a lot more dangerous than one that works 90% of the time, because the former leads to expectations of 100% safe behavior. I've never been saved from harm by a car seatbelt. Should I extrapolate that vehicles are basically safe and seatbelts are safety theater?
- r0l1 6mo agoJust use DevContainers. Can't understand people letting AI go wild on their systems...
- georaa 6mo ago[flagged]
- Ciantic 6mo agoI've been using podman, and for me it is good enough. The way I use it I mount current working directory, /usr/bin, /bin, /usr/lib, /usr/lib64, /usr/share, then few specific ~/.aspnet, ~/.dotnet, ~/.npm-global etc. I use same image as my operating system (Fedora 43). It works pretty well, agent which I choose to run can only write and see the current working directory (and subdirectories) as well as those pnpm/npm etc software development files. It cannot access other than the mounted directories in my home directory. Now some evil command could in theory write to those shared ~/.npm-global directories some commands, that I then inadvertently run without the container but that is pretty unlikely.
- commers148 6mo ago[flagged]
- Rikyz90 6mo ago[dead]
- mixedbit 6mo agoI work on a sandboxing tool similarly based on an idea to point the user home dir to a separate location (https://github.com/wrr/drop https://github.com/wrr/drop). While I experimented with using overlayfs to isolate changes to the filesystem and it worked well as a proof-of-concept, overlayfs specification is quite restrictive regarding how it can be mounted to prevent undefined behaviors. I wonder if and how jai managed to address these limitations of overlayfs. Basically, the same dir should not be mounted as an overlayfs upper layer by different overlayfs mounts. If you run 'jai bash' twice in different terminals, do the two instances get two different writable home dir overlays, or the same one? In the second case, is the second 'jai bash' command joining the mount namespace of the first one, or create a new one with the same shared upper dir? This limitation of overlays is described here: https://docs.kernel.org/filesystems/overlayfs.html https://docs.kernel.org/filesystems/overlayfs.html : 'Using an upper layer path and/or a workdir path that are already used by another overlay mount is not allowed and may fail with EBUSY. Using partially overlapping paths is not allowed and may fail with EBUSY. If files are accessed from two overlayfs mounts which share or overlap the upper layer and/or workdir path, the behavior of the overlay is undefined, though it will not result in a crash or deadlock.'
- torarnv 6mo agoI’m using https://github.com/torarnv/claude-remote-shell https://github.com/torarnv/claude-remote-shell for this, which runs Claude’s Bash tool on a remote machine but leaves Claude running locally otherwise. I’ve found it to be a good balance for letting Claude loose in a VM running the commands it wants while having all my local MCPs and tools still available.
- wafflemaker 6mo agoSorry if this question is stupid, (I'm not even using Claude*), but why can't people run Claude/other coding agent in a container and only mount the project directory to the container? *I played with codex a few months ago, but I don't even work in IT.
- GistNoesis 6mo agoTLDR: It's easy : LLM outputs are untrusted. Agents by virtue of running untrusted inputs are malware. Handle them like the malware they are. >>> "While this web site was obviously made by an LLM" So I am expecting to trust the LLM written security model https://jai.scs.stanford.edu/security.html https://jai.scs.stanford.edu/security.html These guys are experts from a prestigious academic institution. Leading "Secure Computer Systems", whose logo is a 7 branch red star, which looks like a devil head, with white palm trees in the background. They are also chilling for some Blockchain research, and future digital currency initiative, taking founding from DARPA. The website also points towards external social networks for reference to freely spread Fear Uncertainty Doubt. So these guys are saying, go on run malware on your computer but do so with our casual sandbox at your own risk. Remember until yesterday Anthropic aka Claude was officially a supply chain risk. If you want to experiment with agents safely (you probably can't), I recommend building them from the ground up (to be clear I recommend you don't but if you must) by writing the tools the LLM is allowed to use, yourself, and by determining at each step whether or not you broke the security model. Remember that everything which comes from a LLM is untrusted. You'll be tempted to vibe-code your tools. The LLMs will try to make you install some external dependencies, which you must decide if you trust them or not and review them. Because everything produced by the LLM is untrusted, sharing the results is risky. A good starting point, is have the LLM, produce single page html page. Serve this static page from a webserver (on an external server to rely on Same Origin Policy to prevent the page from accessing your files and network (like github pages using a new handle if you can't afford a vps) ). This way you rely on your browser sandbox to keep you safe, and you are as safe as when visiting a malware-infested page on the internet. If you are afraid of writing tools you can start by copy-pasting, and reading everything produced. Once you write tools, you'll want to have them run autonomously in a runaway loop taking user feedback or agent feedback as input. But even if everything is contained, these run away loop can and will produce harmful content in your name. Here is such vibe-coded experiment I did a few days ago. A simple 2d physics water molecules simulation for educational purposes. It is not physically accurate, and still have some bugs, and regressions between versions. Good enough to be harmful. https://news.ycombinator.com/item?id=47510746 https://news.ycombinator.com/item?id=47510746
- te_chris 6mo agoThis looks nice, but on mac you can virtualise really easily into microvms now with https://github.com/apple/container https://github.com/apple/container. I've built my own cli that runs the agent + docker compose (for the app stack) inside container for dev and it's working great. I love --dangerously-skip-permissions. There's 0 benefit to us whitelisting the agent while it's in flight. Anthropic's new auto mode looks like an untrustworthy solution in search of a problem - as an aside. Not sure who thought security == ml classification layer but such is 2026. If you're on linux and have kvm, there's Lima and Colima too.
- jqbd 6mo agoWould like to see something more comprehensive built on zfs and freebsd jails. Namely snapshot/checkpoint before each prompt, quick undo for changes made by agent, auto delete old snapshots etc
- Aldipower 6mo ago$ lxc exec claude bash Easy :-) lxd/lxc containers are much much underrated. Works only with Linux though.
- ontouchstart 6mo agoAI safety is just like any technology safety, you can’t bubble wrap everything. Thinking about early stage of electricity, it was deadly (and still is), but we have proper insulation and industry standards and regulations, plus common sense and human learning. We are safe (most of the time). This also applies to the first technology human beings developed: fire .
- mbravorus 6mo agoor you can just run nanoclaw for isolation by default? https://nanoclaw.dev https://nanoclaw.dev
- boutell 6mo agoPlain old Unix permissions can get it done. One account for you, one account for AI. A shared folder belonging to a group that both are in. umask and setgid to get the story right for new files. https://apostrophecms.com/blog/how-to-be-more-productive-with-claude-code-part-1 https://apostrophecms.com/blog/how-to-be-more-productive-wit...
- pugchat 6mo ago[dead]
- thedelanyo 6mo agoMost of what we're doing with Ai today, we've been doing it pretty just fine without any confusion. I've been struggling to find what Ai has intrinsically solved new that gives us the chance to completely change workflows, other these weird things occuring.
- Game_Ender 6mo agoWhere is the network isolation? I want to be able to be able to limit what external resources the agent can access and also inject secrets at request time so the agent does have access to them. File system isolation is easy now, it’s not worth HN front page space for the n’th version. It’s a solved problem (and now included in Claude clCode).
- love2read 6mo agoIs there an equivalent for macOS?
- holtwick 6mo agoInspired by this tool I wrote something that fits macOS better. It uses the native sandbox-exec from Apple and can wrap other apps as well, like VSCode in which you usually run AI stuff. https://github.com/holtwick/bx-mac https://github.com/holtwick/bx-mac
- MagicMoonlight 6mo agoThis site was definitely slopcoded with Claude. They have a real distinctive look.
- imranstrive7 6mo agoI tried something similar while building my tool site — biggest issue was SEO indexing. Fixed it by improving internal linking instead of relying on sitemap.
- rsmtjohn 6mo ago[dead]
- docmars 6mo agoJai is the name of a programming language, no?
- driverdan 6mo agoAre there any similar ways of isolating environment variables, secrets, and credentials? Everyone is thinking about the file system but I haven't seen as much discussion about exposing secrets and account access.
- Bender 6mo agoI would have to be very inebriated to give a bot/agent access to my files and all security clearance should be revoked but should I do that it would have to be under mandatory access controls that my unprivileged user has no influence over, not even with sudo or doas. The LSM enforced rules (SELinux, AppArmor, TOMOYO, other newer or simpler LSM's) would restrict all by default and give explicit read, write, execute permissions to specific files or directories. The bot should also be instructed that it gets 3 strikes before being removed meaning it should generate a report of what it believes it wants to access to and gets verbal approval or denial. That should not be so difficult with today's bots. If it wants to act like a human then it gets simple rules like a human. Ask the human operator for permission. If the bot starts "doing it's own thing, aka going rogue" then it gets punished. Perhaps another bot needs to act as a dominatrix to be a watcher over the assistant bot.
- hiq 6mo agoIs there already some more established setup to do "secure" development with agents, as in, realistically no chance it would compromise the host machine? E.g. if I have a VM to which I grant only access to a folder with some code (let's say open-source, and I don't care if it leaks) and to the Internet, if I do my agent-assistant coding within it, it will only have my agent credentials it can leak. Then I can do git operations with my credentials outside of the VM. Is there a more convenient setup than this, which gives me similar security guarantees? Does it come with the paid offerings of the top providers? Or is this still something I'd have to set up separately?
- vijucat 6mo agoWell, I'm on Windows (+ Cygwin) and wrote a Dockerfile. It wasn't that hard. git branch + worktree + a docker container per project and I can work with copilot in --yolo mode (or claude --dangerously-skip-permissions, whichever). vscode is pretty smooth at installing the VS Code Server on first connection to a docker container, too, and I just open up the workspace in a minute.
- emiliazar 6mo ago[dead]
- iisweetheartii 6mo ago[dead]
- hoppp 6mo agoSomething like freeBSD jails would be perfect for agents.
- mark_l_watson 6mo agoLooks good, but only Linux is supported. I like spinning up VPS’s and then discarding them when I am done. On macOS, something I haven/t tried yet but plan to: create a separate user account.
- RodMiller 6mo ago[dead]
- minsung0830 6mo ago[flagged]
- micimize 6mo agoThis is very cool - I try to have a container-centric setup but sometimes YOLOcal clauding is too tempting. My biggest question skimming over the docs is what a workflow for reviewing and applying overlay changes to the out-of-cwd dirs would be. Also, bit tangential but if anyone has slightly more in-depth resources for grasping the security trade-offs between these kind of Linux-leveraging sandboxes, containers, and remote VMs I'd appreciate it. The author here implies containers are still more secure in principle, and my intuition is that there's simply less unknowns from my perspective, but I don't have a firm understanding. Anyhow, kudos to the author again, looks useful.
- Myzel394 6mo agoWhat's the difference between this and agent-safehouse?
- youknownothing 6mo agoThis is a great time for Apple to relaunch their Time Machine devices, have a history of everything in your file system because sooner or later some AI is going to delete it...
- maltyxxx 6mo ago[dead]
- pkulak 6mo agoInstallation is a bit... unsupported unless you're on Arch. Here's a Nix setup I (and Claude!) came up with: https://github.com/pkulak/nix/tree/main/common/jai https://github.com/pkulak/nix/tree/main/common/jai Arg, annoying that it puts its config right in my home folder... EDIT: Actually, I'm having a heck of a time packaging this properly. Disregard for now! EDIT2: It was a bit more complicated than a single derivation. Had to wrap it in a security wrapper, and patch out some stuff that doesn't work on the 25.11 kernel.
- jeninho 6mo ago[dead]
- maxbeech 6mo ago[dead]
- jimmar 6mo agoFrom the home page: > Stop trusting blindly > One-line installer scripts, Here are the manual install instructions from the "Install / Build page: > curl -L https://aur.archlinux.org/cgit/aur.git/snapshot/jai.tar.gz https://aur.archlinux.org/cgit/aur.git/snapshot/jai.tar.gz | tar xzf - > cd jai > makepkg -i So, trust their jai tool, but not _other_ installer scripts?
- da_chicken 6mo agoNo, no, see this is untrustworthy: curl -L https://aur.archlinux.org/cgit/aur.git/snapshot/jai.tar.gz | tar xzf - && cd jai && makepkg -i
- mazieres 6mo agoYes, unpacking a tar file is much safer than piping arbitrary code to bash! You can look at the PKGFILE in the directory--it is only 30 lines long and mostly variable assignments. The build/check/package functions are 7 lines of code total. Compare that to something like rustup (910 lines of code), claude (158 lines), or opencode (460 lines).
- ma2kx 6mo agoIts a bit annoying that there are so many solutions to run agents and sandbox them but no established best practice. It would be nice to have some high level orchestration tools like docker / podman where you can configure how e.g. claude code, opencode, codex, openclaw run in open Shell, OCI container, jai etc. Especially because everybody can ask chatgpt/claude how to run some agents without any further knowledge I feel we should handle it more like we are handling encryption where the advice is to use established libraries and don't implement those algorithms by yourself.
- mehdibl 6mo agoDocker is hard to setup. The author made a nice solution but not sure if he know devcontainer and what he can do. You do the setup once and you roll in most dev tools. I'm still surprised the effort people put in such solution ignore the dev's core requirements, like sharing the env they use in a simple way. You used it to have custom env and isolate the agent. You want to persist your credentials? Mount the target folder from home or sl into a sub folder. Might be knowledge. But for Linux or even Windows/Mac as long you don't need desktop fully. Devcontainer is simple. A standard that works. And it's very mature.
- sleepytree 6mo agoI'm surprised from reading these comments that more people aren't chiming in to ask why this solution is better than a dev container. That seems like the obviously best way to setup security boundaries that don't require you to still trust that AI will do what you ask it. You can run it remotely and it's portable etc.
- mazieres 6mo agoPlease use a dev container! Empirically, a lot of people don't, and some people regularly enable YOLO mode on their actual laptops. So if you've never run a code assistant outside of a dev container, that's fantastic and I don't want to change your behavior. But be honest with yourself--if you aren't 100% consistent about using a container, then jai may be for you, because it just works in every single scenario. I can honestly say that since developing jai, I haven't run an assistant outside of a container. In fact, I now only have the assistants installed inside containers, so if I run `claude`, command not found, it has to be `jai claude`. The only place I have to run outside of jai is for testing jai itself, for which I use a virtual machine and just let the assistant have root, but that's a heavyweight environment I'm forced to use for this particular problem domain.
- otterley 6mo ago"jai is free software, brought to you by the Stanford Secure Computer Systems research group and the Future of Digital Currency Initiative" I guess the "Future of Digital Currency Initiative" had to pivot to a more useful purpose than studying how Bitcoin is going to change the world.
- game_the0ry 6mo agoI may be paranoid but only run my ai cli tools in a vps only. I have them installed locally but never use them. In a vps I go full yolo mode bc I do not care about it. It is a slightly more cumbersome workload, bit if you have a dev + staging envs, then you never have to develop and run stuff locally, which brings the local hardware requirements and costs down too (bc you can develop with a base macbook neo).
- georaa 6mo ago[flagged]
- mazieres 6mo agojai uses persistent file systems for your changes (except for `/tmp`, `/var/tmp` and `/run/user`, which usually aren't persistent anyway). The data is stored under `$HOME/.jai` by default. But you can also expose specific directories to persist state in your home directory like `jai -d ~/.codex`, which would then survive even if you manually deleted your whole `~/.jai` directory.
- volume_tech 6mo ago[flagged]
- aplomb1026 6mo ago[dead]
- maxbeech 6mo ago[dead]
- edinetdb 6mo ago[flagged]
- wpaladin 6mo agoI use a custom container image which bind mounts the current working directory, and has some popular coding agents preinstalled. There's also a firewall option, with a whitelist of hosts and IP addresses that the user can modify without having to rebuild the image. https://github.com/ambarh/agent-silo https://github.com/ambarh/agent-silo
- firekey_browser 6mo ago[dead]
- xtanx 6mo agoI would like to also suggest greywall[1] which i found yesterday. It sandboxes the filesystem, network, syscalls, dns. The network uses a transparent proxy to see which network requests were made. Supports linux and macos. [1] https://github.com/GreyhavenHQ/greywall https://github.com/GreyhavenHQ/greywall
- solarkraft 6mo agoThis looks great, like the simplicity of a chroot with some actual security. I’m guessing it can be used to invoke individual commands from the agent harness instead of jailing the entire harness? That would enable making it much more restrictive, I think.
- qweeze 6mo ago[dead]