3 ms·
I was wondering if I should try to create my own pseudo filesystem with FUSE for easy copy-on-write/snapshots, a native feel, and automated secrets filtering/sw
by radio879 16d ago
I was wondering if I should try to create my own pseudo filesystem with FUSE for easy copy-on-write/snapshots, a native feel, and automated secrets filtering/swapping. I might as well combine that with good/easy isolation.
The whole “what sandbox/VM/microVM/thing is best?” question has been bugging me a lot lately, and I no longer trust any of these AI companies to keep data safe.
I’ve been testing a bunch of sandbox-related projects. Sometimes I just use a full Fedora Workstation VM inside Windows 11 with a shared folder, copy a project into it, and run long agent tasks there. It’s not ideal, but it is pretty safe. Sometimes I run agents in different WSL2 distros and test different things inside those.
I did like gVisor from Google — it’s not quite a microVM, but it’s not really just a normal container either. It wasn't easy to figure out how to get it working tho. Lima Machines worked well too, and I don’t remember it being annoying. SmolVM... ugh. There are two projects with exactly the same name, and it got confusing enough that I gave up. One of them did work when I tried it, though.
The confusing part is that there are now hundreds of sandbox projects, and they all solve slightly different pieces of the problem. Some have filesystem isolation, some have networking controls, some handle credentials better, etc. Nono, for example, has a nice secrets filtering/swapping idea where real credentials can be replaced with dummy values, but there have also been GitHub reports about isolation gaps — data being accessible when it isn’t supposed to be. I’m trying to figure out which projects are worth using, which are worth skipping entirely, and which might just have useful pieces of code or ideas to borrow.
I’ve got GPT-5.6 in one window doing a fairly ridiculous analysis of the different approaches and the likely long-term reliability/adoption risk of each repo. Separately, I have a WSL2 distro running Reasonix with DeepSeek doing its own analysis so I can compare conclusions.
What I eventually want is a desktop GUI over whatever combination of sandbox technologies turns out to be reliable. Ideally I could just type:
“Spin up 5 sandboxes for project X. Put Claude Code in one, Reasonix in #2, Codex in #3…”
or:
“Create 3 sandboxes, put whatever coding agents in 1, 2, and 3, and then have each one run twice.”
If it’s AI-powered, it could automatically name folders and copy results back somewhere like `folderName_3a`, or use Git branches/worktrees if desired. I don’t always want to use Git.
Every sandbox CLI has its own syntax, code quality, reliability, ease/pain of getting it working, configuration format, mount rules, networking options, etc., and I don’t particularly enjoy memorizing another pile of commands just to isolate an agent.
I’ve tried quite a few of them. A lot of them are still rough enough that I hit errors quickly and move on. Some seem much more mature — Lima is one I like conceptually, although native Windows support would be nice but I guess not a huge deal.
Credentials are something I never cared much about (API keys and stuff like that) but now.... I'm more worried. I really don’t want to deal with any problems from that. Or something installing something that grabs SSH keys, browser passwords (FYI.. Z Code asks you "do you wanna import all the logins from chrome?) browser sessions, cloud credentials, or my whole home directory. That concern isn’t limited to Chinese software either. I don’t automatically trust US AI companies just because they’re US companies. Zuck, Elon...zero trust in those two.
So I’m increasingly thinking the “right” answer might not be one sandbox project at all. It may be a GUI/orchestration layer that combines more than one backend and more than one type of sandbox. There could be common default presets and combinations of Git worktrees plus containers and/or VMs. I also feel safer that Docker/Podman on Windows generally runs inside WSL2, because it’s basically containers inside a VM.
The goal would be strong isolation underneath — maybe even combining two or more layers so one failure doesn’t expose everything — plus explicit project-folder mounts with read-only or read/write options, rollback/snapshots, network controls, secrets substitution, disposable environments, and an easy way to fan the same task out to multiple agents/models.
I also like the idea of having an AI model in front of the whole thing, with the ability to save whatever setup it creates as a preset so the AI part can be skipped next time. And I want it to support not only parallel agents using different models, but also loops where the exact same agent setup runs several times.