5 ms·
The login is annoying but, lacking an open source alternative, this has been my daily driver for a while now because it works great out of the box with two key
by rusch 2mo ago
The login is annoying but, lacking an open source alternative, this has been my daily driver for a while now because it works great out of the box with two key features: outbound firewall and secret injection with placeholders.
I run it with superset and then each git worktree is mounted in a sandbox that is configured for each repo i work in.
Closest open source I have seen is https://earendil-works.github.io/gondolin https://earendil-works.github.io/gondolin but the DX is not as polished. https://exe.dev/ https://exe.dev/ would be perfect but it does not come with outbound firewall.
Does anyone have a better alternative?
- SegmentTree 2mo agoEclipse Enclave does exactly that: There is an outbound firewall and secret injections, so that the agent never sees a real key. And it's fully open source: https://github.com/eclipse-enclave/enclave https://github.com/eclipse-enclave/enclave
- deleted 2mo ago[deleted]
- _ink_ 2mo agoDoes secret injection really prevent that the agent send my GitHub key somewhere? If it has access to it via env var, can it not just paste it somewhere?
- rusch 2mo agoThe env var is just a placeholder in the VM, so no real secret is in there.
- llimllib 2mo agoright, but say you give the agent access to github and it can push as you, or make a gist; now it can easily exfiltrate your secret. And that's just an easy case - really if it has any network access at all it can come up with a clever way to route a request through the network such that the key comes back somewhere in the request. If you scan for it inbound too, the machine can obfuscate it. Our agents are trained to be so intensely helpful and they have such intricate knowledge of how things work that they will do some incredibly clever tricks to do what you ask them to do.
- skinfaxi 2mo agoThe agent has no access to the secret. It has a placeholder that is replaced at a higher level. When it makes the network request the secret is substituted but that is outside of the caller's worldview.
- Eldt 2mo agoSo what stops it from sending a network request to a git repo that pushes what that placeholder resolves to?
- skinfaxi 2mo agoHow would that work? You don't control github.com servers so your repo would never see the secret. edit: You may want to look into tokenizing proxies as the general application of this concept.
- llimllib 2mo agoYour agent writes secret.txt with the placeholder, and the tokenizing proxy replaces it with the token, then the agent reads secret.txt
- throwaw12 2mo agoI havent used nor gondolin neither docker's solution, but curious to know what gondolin is missing (evaluating both for my personal use)? is it only the DX or something else, if DX, can you what exactly is missing? thanks
- rusch 2mo agoYes, the stated "target workload"[0] is not what i'm looking for. I want my agent to run for long, spin up dedicated local stack while developing etc. It seems with gondoling i need to explain the agent to run commands in the sandbox, but then where does the agent run itself? [0]: https://earendil-works.github.io/gondolin/workloads/ https://earendil-works.github.io/gondolin/workloads/
- eli 2mo agoYou can run the agent in the gondolin sandbox if you wish. Their example implementation with pi uses a pi extension so that pi runs on the host but the read/write/bash/etc tools run in the guest. Doesn’t have to be that way though.
- olejorgenb 2mo agoNote that the pi-extension example in the gondolin repo is very outdated. Look in the pi repo instead.
- codethief 2mo agoIn my experience it's mostly the UX/DX where Gondolin is lacking. For instance, I don't want to set up a JavaScript project every single time I need a sandbox. Instead, I just want to place a config file somewhere in my repo or my home dir and be done with it. So I wrote a wrapper around Gondolin which allows me to do that and a few other things: https://github.com/codethief/tuor https://github.com/codethief/tuor (Warning: Still very much experimental / underdocumented.)
- dbmikus 2mo ago
- sparsesignal 2mo agoWhat I run is one hardened QEMU/KVM VM per project holding the whole dev environment (editors, agents, containers), with nftables on the host allowing internet egress but dropping anything aimed at the host, the LAN, or any other private address, plus an allowlist for deliberate exceptions. Basically, it's a plain QEMU/KVM VM on a stock Debian cloud image: device model stripped down to a virtio disk, a virtio NIC and a serial console, nested virt off, no passwordless sudo in the guest. It also ships a containment check that scans outward from inside the guest, so the network boundary is something you can verify. Wrapping the whole environment rather than a single agent session puts supply chain attacks inside the boundary too. A poisoned npm or PyPI package, or a compromised editor extension, lands in the VM instead of on the host. That was the original reason I set this up; agents just made it more urgent. There's no per-domain egress allowlist; the policy is "internet yes, private addresses no". Secret injection isn't built in either, though Infisical's agent-vault on the host as an egress proxy covers that part. Wrote the whole setup up here, in case it's useful: https://karamatli.com/posts/network-isolated-kvm-sandbox-ai-agents/ https://karamatli.com/posts/network-isolated-kvm-sandbox-ai-...
- someothherguyy 2mo agoI currently do something similar, but this article was a nice read and gave me some new ideas.
- TacticalCoder 2mo agoYeah I discovered your blog a few days ago: I've got a setup not unlike yours. > So rather than pick one, this post advocates layering both, in the spirit of defense in depth: a sandbox VM wraps your containers along with the whole toolchain, and that sandbox reaches the internet but has no route to anything private. Yup it's the only proper way. And that is true not just for AI harnesses/agents (that shall try to escape), but also for stuff like Plex/Jellyfin/Immich/private pastebin etc. If you care about security, there really simply is zero reason to run containers on the bare metal.
- binsquare 2mo agomaintainer, I would recommend trying out: https://github.com/smol-machines/smolvm https://github.com/smol-machines/smolvm It has network filtering + placeholders for secrets. OSS, no logins needed
- reddec 2mo agoI've put some effort to integrate it to my agentic workflow. The problem, however, with docker in smolvm: it work-ish (there is example), but quite hacky. Another problem which I wasnt able to solve - persistent image without Dockerfile. CloudInit will be ideal. Documention at this moment in an early stage. Overall, its a great project but for me was simpler just use Virtual Machine Manager (libvirt GUI). I wish all luck to the maintainers, but probably DX-wise I will prefer to have more granular or predictable controls (eg micro cloud from Canonical).
- icedchai 2mo agoThere's also microsandbox which has similar features: https://github.com/superradcompany/microsandbox https://github.com/superradcompany/microsandbox (Not affiliated with them, just tried it out last week.)
- binsquare 2mo agoI'm aware of them! Yep - similar in some ways but headed towards different directions. I am building a virtual machine to simplify/replace container infra. Ex. we run containers inside of linux VM's even in the `cloud`, resulting in managing both the vm, and the containers. But smol machines is a lightweight, portable VM that you can package into a single portable .smolmachine file to be rehydrated on any platform, kind of like how containers are used for today. Sandboxing happens to be a feature of virtual machines, so we are alike in being used for sandboxing.
- jachris 2mo agoAgreed. Network control and secret injection together with a microVM setup is as good as it gets right now, although I believe that we need more fine-grained tools down the road. It sounds like Microsandbox would be the perfect fit for what you are describing. I also built my own coding agent workbench on top of it (https://github.com/isolade/isolade https://github.com/isolade/isolade). Microsandbox is quite cool, check it out: https://github.com/superradcompany/microsandbox https://github.com/superradcompany/microsandbox
- MikhailTal 2mo agohttps://github.com/e2b-dev/e2b https://github.com/e2b-dev/e2b
- devttyeu 2mo agoFor persistent-ish one-off apps https://xbin.dev/ https://xbin.dev/ (my project) Has egress and ingress filtering, egress can be bound to host/internet/subnet or even better to internal apps (which are each separate netns) meaning you can do your own firewall/vpn/whatever per sandbox. Plus you control what other components in the sandbox env the app can communicate with. Really not built for day-to-day dev work though, more like automating your company/life / getting rid of SaaS (e.g. for technical Founders / Sales etc, not exactly useful for dev work)
- petesergeant 2mo agoWhat specifically do you want? I have: https://github.com/pjlsergeant/byre https://github.com/pjlsergeant/byre -- slightly different security model, but lazer-focused on developer experience; my daily driver and I love it not just because I wrote it. The TUI is great for configuring and setting up instant boxes just how you want https://pleasedonotescape.com/ https://pleasedonotescape.com/ -- a list of every other agent jail I could find, filterable by open-source and whatever else you want
- nzjrs 2mo agoIf you just need a python+venv sandbox with dev-first UX, no container build step needed, and no startup cost then I am using https://github.com/nzjrs/sandbubble https://github.com/nzjrs/sandbubble in prod.
- mellowagain 2mo agoive been using https://github.com/Gerharddc/litterbox https://github.com/Gerharddc/litterbox
- rpoisel 2mo agoDon't want to say it's better, but I implemented Agent Circus (https://github.com/Embedded-Focus/agent-circus https://github.com/Embedded-Focus/agent-circus) which allows to lock AI agent harnesses into docker containers. I'm using it as my main driver since months. Support for running agent harnesses in unprivileged podman containers is on my feature list. :-)
- pojzon 2mo agoDocker containers are not enough isolation for anyone that cares about jailbreak scenarios. Only real alternative is to use microvms. My goto solution for this are apple/containers.
- Supermancho 2mo ago> Docker containers are not enough isolation for anyone that cares about jailbreak scenarios. For the vast majority of developers, containers are enough, which is why they are ubiquitous while vms are less common. Ofc that ubiquity has led to lazy configuration, which is how the jailbreaking can occur. Knowing what you are doing with containers is a requirement to use containers as an AI sandbox.
- wbl 2mo agoThe AI launches new kernel bugs as matter of course.
- zmmmmm 2mo agothe big thing containers don't allow is for the agent to run and use docker itself without compromising the host I'm not sure where "vast majority" cuts in but I would say a huge number of developers use docker and it is inconvenient at best if your AI harness can't actually run and test the infra it is building against
- pjmlp 2mo agoThere were not the solution for a while now, that is why Kata containers came to be in first place.
- pbasista 2mo agoI use Linux Containers managed by Incus for working with Claude. I have a dedicated container for that. It can run its own Docker daemon and other system services if needed. Apart from the Claude login token, it has no SSH keys or other credentials. I push everything I need to it from the local machine. And I pull the Claude generated outputs from it. Of course, this kind of setup requires a stack which can run or at least be tested without any credentials.
- Tomte 2mo agoI do the same, but with pi.dev in Incus, mapping a project folder into the VM. What I don‘t have compared to sbx is an outbound firewall, but my VM does not have any personal/interesting data, only a vanilla Fedora installation and the project dir with open source code, so I do not care much about exfiltration.
- radlad 2mo agoLove Incus and I'm using throwaway restricted projects for testing. Highly recommend incus-windows if you need to do any Windows testing. Having agents validate Windows behavior has reduced so much toil for me.
- indigodaddy 2mo agoHere's my incus thing meant for a VPS/server: https://GitHub.com/jgbrwn/vibebin https://GitHub.com/jgbrwn/vibebin
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- PufPufPuf 2mo agoI made Locki: https://github.com/JanPokorny/locki https://github.com/JanPokorny/locki Internally uses a single VM + Incus containers, supports docker/Kubernetes in each sandbox, has integrated worktree management.
- 384028345 2mo agoIsn't Nvidia's openshell exactly what you're looking for? I'm asking because I'm just learning about this stuff myself and tested openshell yesterday with pi for the first time. https://github.com/NVIDIA/openshell https://github.com/NVIDIA/openshell
- intrasight 2mo agoI have a colleague who's using openshell. The advantage is it's independent of the containerization layer, yes?
- cpburns2009 2mo agoGondolin looks interesting. It sounds like a TypeScript wrapper that achieves the same thing as my setup: Docker & Kata Containers 4 (KVM/QEMU backend) for microVMs, iron-proxy for egress and secrets, and dnsmasq for internal network name resolution (workaround for a Docker/Kata incompatibility).
- olejorgenb 2mo agoI'd say it's not ready for prime time yet unfortunately: https://github.com/earendil-works/gondolin/issues/115 https://github.com/earendil-works/gondolin/issues/115
- sylvinus 2mo agoAre we all posting our agent VM containers ? :) https://github.com/sylvinus/agent-vm https://github.com/sylvinus/agent-vm
- Thanemate 2mo agoWe're basically at the point where people can build their own "X", with "X" being internal tooling.
- olejorgenb 2mo agoYes, but for certain things like sandboxing I hope we can converge on a handful robust, well-tested/audited solutions...
- esafak 2mo agoWhat is different about yours?
- kstenerud 2mo agoI wrote one that has both of those: yoloAI (MIT, Go, single binary, no login). https://github.com/kstenerud/yoloai https://github.com/kstenerud/yoloai Outbound firewall is `--network-isolated`: egress is denied except the agent's own API endpoints plus domains you allow, enforced sandbox-side (working on host-side enforcement now). `--network-none` if you want nothing. Credential brokering works the way you describe (currently Claude-only, I'll add more as time allows). The API key stays on the host, a local proxy injects it into the outbound request, and the sandbox never holds anything worth stealing. Other agents' credentials currently arrive as read-only file mounts instead (weaker, and something I'll fix soon). Generalising the injector is the obvious next thing. One difference from your setup: yoloAI copies your worktree instead of mounting it. The agent works on the copy, you `yoloai diff`, and `yoloai apply` replays the commits into your real repo. That's deliberate. Docker's own security docs talk about the dangers of bombs being left behind in a live-mounted dir (git hooks, package.json scripts, Makefiles, IDE task config), which diff/apply avoids. Isolation is per-sandbox rather than fixed: runc, gVisor, or Kata VMs (QEMU or Firecracker) on Linux; Seatbelt or full macOS VMs via Tart on a Mac.
- westurner 2mo ago> yoloAI copies your worktree instead of mounting it Cloudflare/artifact-fs does lazy shallow git clones with a FUSE filesystem. https://github.com/cloudflare/artifact-fs https://github.com/cloudflare/artifact-fs Would that be faster? Re: sandboxing methods like Clawk, Amla sandbox, bwrap, agentvm, ARM64 MTE with wasmtime-mte: https://news.ycombinator.com/item?id=48893850 https://news.ycombinator.com/item?id=48893850 A few months ago now I started adding seccomp sandboxing to jinja2rs and then liboverlayfs support to ansiblers (which are early Rust ports). Haven't finished that, but I started working on a VM format that stores signed machine state into an OCI container repository, using the hypervisor migration support of KVM/QEMU. Though this is not safe yet if ever, VM migrations are probably another way to sandbox and deploy en masse.
- kstenerud 2mo ago[dead]
- 2mo ago
- taude 2mo agowhat about coder (and coder workspaces)? https://github.com/coder/coder https://github.com/coder/coder
- chrisweekly 2mo agohttps://smolmachines.com https://smolmachines.com has "smolvm" microvms, for better security. The DX is whatever you decide to do with it.
- dbmikus 2mo agoMy startup is open core: https://github.com/gofixpoint/amika https://github.com/gofixpoint/amika We run cloud sandboxes, and have some experimental local sandbox support that is fully OSS. Main thing for amika.dev is you can control the sandboxes and agents interchangeabley by SSH, web, or API, and can expose the services the agent is working on over signed URLs Entire sandbox config is a TOML file We're going to improve the local OSS sandbox mode and add better network controls over the next couple weeks. Ultimately, what we're building kind of like if Tailscale and Firecracker had a baby, with a messaging protocol for remote controlling any sandboxed agent It's free to try out. Still a lot to build, so we really appreciate any and all feedback about what we should focus on
- zippie 2mo agoWe run https://github.com/NVIDIA/OpenShell https://github.com/NVIDIA/OpenShell on our workloads and we like it because it is backed by NVIDIA and fits natively within our k8s env. It also provides a nice TUI for network policy management and prebuilt sandboxes which have claude code, codex, etc.
- notsirius 2mo agoOoh - can you share more about your setup with superset? I tried getting it integrated with superset a while ago with no dice.
- rusch 2mo agoCurrently running codex. I run one sandbox per repo. So I create the sandbox in the repo root. Then it's a custom terminal preset: sbx run --name "yoursandbox" -- --cd "$PWD" This boots a sbx session in the worktree directory. For Claude there is no --cd so it's more hacky, but I solved it by creating a sbx kit with entrypoint script that reads a flag (e.g --cwd) from the terminal preset command and then inside the sandbox cd's there and starts claude.
- notsirius 2mo agoAh - so youre running in direct mode then? https://docs.docker.com/ai/sandboxes/workflows/#direct-mode https://docs.docker.com/ai/sandboxes/workflows/#direct-mode Def makes integrating easier but I try to avoid using direct mode for security (exposes .git folder, though there's probs a better way to protect it by disabling hooks or something)
- rusch 2mo agoYes correct. We don't use git hooks so they are globally disabled. I also see they added host worktree mode which could work well with superset since it creates the worktrees. https://docs.docker.com/ai/sandboxes/workflows/#host-worktree https://docs.docker.com/ai/sandboxes/workflows/#host-worktre...
- hanwenn 2mo agoI wrote my own (or rather, had Claude write it), https://github.com/hanwen/runclaude https://github.com/hanwen/runclaude. Mine uses containers, and makes only the git/jj workspace read/write, hiding all credentials that are in my home dir.
- schmitthub 2mo ago[dead]
- isityettime 2mo agoThis has both features and is open-source: https://nono.sh/ https://nono.sh/ It's an OS-level sandbox, though. It doesn't launch VMs or containers for sandboxing purposes; it uses whatever sandboxing features your host kernel offers.
- grudg3 2mo agoWouldn't say 'better' alternative, but I worked on making my own setup that I can trust by implementing a pi extension that leverages smolvm and agent-vault. The VM tooling is controlled by nix flakes. I can't share the source code (developed on company time), but I have a 'spec' of the whole thing, which you should be able to feed to your agent to replicate - https://gist.github.com/mahalel/c4e984292ff90bd4e1126955515815cd https://gist.github.com/mahalel/c4e984292ff90bd4e11269555158...
- appcypher 2mo agoit's been mentioned on this thread already, we think https://github.com/superradcompany/microsandbox/ https://github.com/superradcompany/microsandbox/ is the closest OSS to it and we have a great DX and improving.
- mplemay 2mo agoIf you are looking to write have you're agent write typescript code, check out: https://github.com/mplemay/belgie https://github.com/mplemay/belgie. TLDR: You're agent will get a isolated v8 runtime (chrome's sandboxed javascript runtime)
- codethief 2mo ago> Does anyone have a better alternative? Not necessarily better but OpenSandbox[0] by Alibaba seems similar. [0]: https://github.com/alibaba/OpenSandbox https://github.com/alibaba/OpenSandbox