9 ms·
Launch HN: Freestyle – Sandboxes for Coding Agents
We’re Ben and Jacob, cofounders of Freestyle (https://freestyle.sh https://freestyle.sh). We’re building a cloud for Coding Agents.
For the first generation of agents it looked like workflows with minimal tools. 2 years ago we published a package to let AI work in SQL, at that time GPT-4 could write simple scripts. Soon after the first AI App Builders started using AI to make whole websites; we supported that with a serverless deploy system.
But the current generation is going much further, instead of minimal tools and basic serverless apps AI can utilize the full power of a computer (“sandbox”). We’re building sandboxes that are interchangeable with EC2s from your agents perspective, with bonus features:
1. We’ve figured out how to fork a sandbox horizontally without more than a 400ms pause in it. That's not forking the filesystem, we mean forking the whole memory of it. If you’re half way down a browser page with animations running, they’ll be in the same place in all the forks. If you’re running a minecraft server every block and player will be in the same place on the forks. If you’re running a local environment and an error comes up in process that error will be there in all the forks. This works for snapshotting as well, you can save your place and come back weeks later.
2. Our sandboxes start in ~500ms.
Demo: https://www.loom.com/share/8b3d294d515442f296aecde1f42f5524 https://www.loom.com/share/8b3d294d515442f296aecde1f42f5524
Compared with other sandboxes, our goal is to be the most powerful. We support full Linux + hardware-virtualization, eBPF, Fuse, etc. We run full Debian with multiple users and we use a systemd init instead of runc. Whatever your AI expects to work on debian should work on these vms, and if it doesn’t send a bug report.
In order to make this possible, we’ve moved to our own bare metal racks. Early in our testing we realized that moving VMs across cloud nodes would not have acceptable performance properties. We asked Google Cloud and AWS for a quote on their bare metal nodes and found that the monthly cost was equivalent to the total cost of the hardware so we did that.
Our goal is to build the necessary infrastructure to replicate the human devloop on the massively multi-tenant scale of AI, so these VMs should be as powerful as the ones you’re used to, while also being available to provision in seconds.
- maxmaio 6mo agoCongrats Ben and Jacob!
- n2d4 6mo agoCool! I've been using your API for running sandboxed JS. Nice to see you also support VMs now. > we mean forking the whole memory of it How does this work? Are you copying the entire snapshot, or is this something fancy like copy-on-write memory? If it's the former, doesn't the fork time depend on the size of the machine?
- benswerd 6mo agoWe're using copy on write with the memory itself. Fork time is completely decoupled from the size of the machine. Creating snapshots takes a 2-4 second interruption in the VM due to sheer IO that we didn't want here. Whats especially cool about this approach is not only is fork time O(1) with respect to machine size, but its also O(1) with respect to the amount of forks.
- _jayhack_ 6mo agoWould love to understand how you compare to other providers like Modal, Daytona, Blaxel, E2B and Vercel. I think most other agent builders will have the same question. Can you provide a feature/performance comparison matrix to make this easier?
- benswerd 6mo agoI'm working on an article deep diving into the differences between all of us. I think the goal of Freestyle is to be the most powerful and most EC2 like of the bunch. Daytona runs on Sysbox (https://github.com/nestybox/sysbox https://github.com/nestybox/sysbox) which is VM-like but when you run low level things it has issues. Modal is the only provider with GPU support. I haven't played around with Blaxel personally yet. E2B/Vercel are both great hardware virtualized "sandboxes" Freestyle VMS are built based on the feedback our users gave us that things they expected to be able to do on existing sandboxes didn't work. A good example here is Freestyle is the only provider of the above (haven't tested blaxel) that gives users access to the boot disk, or the ability to reboot a VM.
- tomComb 6mo agoAnd fly.io sprites
- benswerd 6mo agoFly.io sprites is the most similar to us of the bunch. They do hardware virtualization as well, have comparable start times and are full Linux. What we call snapshots they call checkpoints. The big pros of Sprites over us is their advanced networking stack and the Fly.io ecosystem. The big cons are that Sprites are incredibly bare bones — they don't have any templating utilities. I've also heard that Sprites sometimes become unavailable for extended periods of time. The big pros of Freestyle over Sprites is fork, advanced templating, and IMO a better debugging experience because of our structure.
- knowsuchagency 6mo agoThanks for the thoughtful response. I'm predominantly a self-hoster, but I think your product makes a lot of sense for a wide variety of users and businesses. I'm excited to try out freestyle!
- Fraaaank 6mo agoYour pricing page is broken
- benswerd 6mo agoReviewing this now. our public pricing at www.freestyle.sh/pricing seems to be working, can you point me in a more specific direction?
- aplomb1026 6mo ago[dead]
- borakostem 6mo ago[flagged]
- benswerd 6mo agoSo this is an ongoing optimization point, no perfect solution exists. Freestyle VMs work with a network namespace and virtual ethernet cable going into them, so they all think they are the same IP. This means that while complex protocol connections like remote Postgres can break in the forks, stuff like Websockets just automatically reconnects.
- MarcelinoGMX3C 6mo agoThe technical challenges in getting memory forking to deliver those sub-second start and fork times are significant. I've seen the pain of trying to achieve that level of state transfer and rapid provisioning. While "EC2-like" gets the point across for many, going bare metal reveals the practical limits of cloud virtualization for high-performance, complex workloads like these. It shows a real understanding of where cloud abstraction helps and where it just adds overhead. The cost argument for owning the hardware for this specific use case also makes sense, considering the scale these agent environments will demand. Also worth noting, sandboxes are effectively an open attack surface; architecting them not to be in your main VPC is a sound security decision from the start.
- skybrian 6mo agoIt doesn't seem very easy to calculate how much it would cost per month to keep a mostly-idle VM running (for example, with a personal web app). The $20/month plan from exe.dev seems more hobbyist-friendly for that. Maybe that's not the intended use, though?
- benswerd 6mo agoWe're not going after hobbyists. We're building the platform for companies like exe.dev to build on. Thats why its all usage based. That said, our $50 a month plan can be used as an individual for your coding agents, but I wouldn't recommend it.
- indigodaddy 6mo agoOoof, if you are the middleman platform then it's sure gonna get expensive for the end user
- rvz 6mo ago> The $20/month plan from exe.dev seems more hobbyist-friendly for that. Maybe that's not the intended use, though? And you can go even below that by self-hosting it yourself with a very cheap Hetzner box for $2 or $5.
- skybrian 6mo agoCan you start up multiple VM's easily on a Hetzner box?
- siva7 6mo agoI have so many interesting problems on Ai, sandboxing isn't one of them. It's a pointless excercise yet disproportionately so many people love to to do this. Probably because sandboxing doesn't feel as magic as Agents itself and more like the old times of "traditional" software development.
- iterateoften 6mo agoYeah, idk I guess it’s interesting if you are an engineer looking for something to do, But like I see multiple sandbox for agents products a week. Way too saturated of a market
- benswerd 6mo agoI disagree (as a sandboxing company). With respect to the market, every single sandbox sucks. I'm not gonna shit talk competitors but there is not a good sandboxing platform out there yet — including me — compared to where we'll be in 6 months. We've heard all the platforms have consistent uptime, feature completeness, networking and debugging issues. And in our own platform we're not 1/10ths of the way through solving the requests we've gotten. Next generation of Agents needs computers, and those computers are gonna look really different than "sandboxes" do today.
- hobofan 6mo agoIt is a mostly pointless exercise if the goal is trying to contain negative impact of AI agents (e.g. OpenClaw). It is a very necessary building block for many common features that can be steered in a more deterministic way, e.g. "code interpreter" feature for data analysis or file creation like commonly seen in chat web UIs.
- rasengan 6mo agoInteresting! We're working on a similar solution at UnixShells.com [1]. We built a VMM that forks, and boots, in < 20ms and is live, serving customers! We have a lot of great tools available, via MIT, on our github repo [2] as well! [1] https://unixshells.com https://unixshells.com [2] https://github.com/unixshells https://github.com/unixshells
- tomComb 6mo agoCan your service scale ram? like the way docker desktop does. Manual is fine.
- benswerd 6mo agoyep you can choose ram + disk + cpu size
- tomComb 6mo ago? You say 'yes' but you seem to be answering a different question. Docker desktop only makes me choose a max ram - it dynamically scales RAM usage. I don't need fully automatic like that, but the ability to vertically scale RAM for an existing instance is really important, particularly given the cost of RAM these days.
- benswerd 6mo agoAh we cannot do this without a restart. Hot pluggable ram is something I'm interested in but is currently a backburner feature.
- johnwhitman 6mo ago[dead]
- stingraycharles 6mo agoI’m super interested since it seems like you have given everything a lot of thought and effort but I am not sure I understand it. When I’m thinking of sandboxes, I’m thinking of isolated execution environments. What does forking sandboxes bring me? What do your sandboxes in general bring me? Please take this in the best possible way: I’m missing a use case example that’s not abstract and/or small. What’s the end goal here(
- benswerd 6mo agoSo isolation is correct. Forking a sandbox gives you multiple exact duplicates of isolated environments. When your coding agent has 10 ideas for what to do, to evaluate them correctly it needs to be able to evaluate them in isolation. If you're building a website testing agent and halfway down a website, with a form half filled out a session ongoing, etc and it realizes it wants to test 2 things in isolation, forking is the only way. We also envision this powering the next generation of devcycles "AI Agent, go try these 10 things and tell me which works best". AI forks the environment 10 times, gets 10 exact copies, does the thing in each of them, evaluates it, then takes the best option.
- indigodaddy 6mo agoYep I can see this especially when the agent is spinning up test servers/smokes and you don't want those conflicting. How do we reconcile all the potential different git hashes though, upstream I guess etc (this might be an easy answer and I'm not super proficient with git so forgive)
- benswerd 6mo agoSo we recommend branch per fork, merge what you like. You have to change the branch on each fork individually currently and thats unlikely to change in the short term due to the complexity of git internals, but its not that hard to do yourself `git checkout -b fork-{whateverDiscriminator}`
- 6mo ago
- stocktech 6mo agoI built something like this at work using plain Docker images. Can you help me understand your value prop a little better? The memory forking seems like a cool technical achievement, but I don't understand how it benefits me as a user. If I'm delegating the whole thing to the AI anyway, I care more about deterministic builds so that the AI can tackle the problem.
- benswerd 6mo agoSo first MicroVM != Container, and container is not a secure isolation system. I would not run untrusted containers on your nodes without extra hardening. The memory forking was originally invented because for AI App Builders and first response driven applications its extremely important that they are instant (difference between running bun dev and the dev server already being running). However its much more generally applicable, Postgres is a great example of this. You can't fork the filesystem under postgres and get consistency. Same thing with a browser state, a weird server state, or anything that exists in memory. The memory forking gives a huge performance boost while snapshotting whats actually going on at one instant.
- sabedevops 6mo agoWhat does this protect you from that you’re exposed to by running a well-crafted rootless container on a system with SELinux or similar?
- benswerd 6mo agoGenerally kernel level attacks and neighbor performance impacts on the security side. On the functional side without a kernel per guest you can't allow kernel access for stuff like eBPF, networking, nested virtualization and lots of important features. Here is a good blog from docker explaining how even the best container is not as safe as a MicroVM https://www.docker.com/blog/containers-are-not-vms/ https://www.docker.com/blog/containers-are-not-vms/ theoretically you can get to fairly complete security via containers + a gVisor setup but at the expense of a ton of syscall performance and disabling lots of features (which is a 100% valid approach for many usecases).
- jnstrdm05 6mo agohow many seconds to provision are we talking about here? 1 sec vs 60 is a dealbreaker for me, some clarity on that would be nice.
- benswerd 6mo ago500ms. Less than 1 second. We're aiming to get that down to 200ms in the next 3 months.
- n1tro_lab 6mo ago[dead]
- vimota 6mo agoThis is awesome - the snapshotting especially is critical for long running agents. Since we run agents in a durable execution harness (similar to Temporal / DBOS) we needed a sandboxing approach that would snapshot the state after every execution in order to be able to restore and replay on any failure. We ended up creating localsandbox [0] with that in mind by using AgentFS for filesystem snapshotting, but our solution is meant for a different use case than Freestyle - simpler FS + code execution for agents all done locally. Since we're not running a full OS it's much less capable but also simpler for lots of use cases where we want the agent execution to happen locally. The ability to fork is really interesting - the main use case I could imagine is for conversations that the user forks or parallel sub-agents. Have you seen other use cases? [0] https://github.com/coplane/localsandbox https://github.com/coplane/localsandbox
- benswerd 6mo agoDeterministic testing of edge cases. It can be really hard to recreate weird edge cases of running services, but if you can create them we can snapshot them exactly as they are.
- dominotw 6mo agodumb question. none of these protect your from prompt injection. yes?
- benswerd 6mo agono, but the goal of these is if you are faced with prompt injection the worst case scenario is the AI uses that computer badly.
- dominotw 6mo agounless i am misundestanding. not sure how this computer prevents secrets from my gmail leaking. thats the worst case.
- benswerd 6mo agoIf you put your gmail credentials into a VM that an AI Agent dealing with untrusted prompts has access to they should be treated as leaked and be disabled immediately. However, if you don't put your administrative credentials inside of the VM and treat it as an unsafe environment you can safely give it minimal permissions to access specific things that it needs and using that access it can perform complex tasks.
- dominotw 6mo agoi am talking about this . not my gmail credentials. https://simonwillison.net/2024/Mar/5/prompt-injection-jailbreaking/#vendor-jailbreaking https://simonwillison.net/2024/Mar/5/prompt-injection-jailbr...
- edf13 6mo ago[dead]
- schopra909 6mo agoHonestly never considered the forking use case; but it makes a ton of sense when explained Congrats on the launch. This is cool tech
- benatkin 6mo agoIt's hard to tell what this is or how it compares to other things that are out there, but what I latched onto is this: > Freestyle is the only sandbox provider with built-in multi-tenant git hosting — create thousands of repos via API and pair them directly with sandboxes for seamless code management. On top of that, Freestyle VMs are full Linux virtual machines with nested virtualization, systemd, and a complete networking stack, not containers. It makes me think of the git automation around rigs in Gas Town: https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04 https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16d... Edit: I realize the Loom is a way to look at it. Loom interrupted me twice and I almost skipped it. However it gave me a better idea of what it does, it "invents" snapshotting and restoring of VMs in a way that appears faster. That actually makes sense and I know it isn't that hard to do with how VMs work and that it greatly benefits from having only part of the VM writable and having little memory used (maybe it has read-only memory too?).
- benswerd 6mo agoSo the snapshotting tech is actually 100% independent of Git. Git is useful for branching vs forking (IE you can't merge two VM forks back together), but all the tech I showed in the Loom exists independently from Git. The hard part of it was making the VM large and powerful while making snapshotting/forking instant, which required a lot of custom VMM work.
- benatkin 6mo ago> The hard part of it was making the VM large and powerful while making snapshotting/forking instant, which required a lot of custom VMM work. I don't find "large and powerful" in reference to a VM to sound compelling. What should be large? The memory? The root disk? As I alluded to in my comment, I'm more curious about what can be made small. Also I'm skeptical that if I forked a vm running a busy Gas Town that it would be very light or fast in how it forks. A well behaved sqlite I could see, but then I'd wonder why not just fork the storage volume containing the database...
- benswerd 6mo ago
- skybrian 6mo agoAny ideas for locking down remote access from an untrusted VM? Cloudflare has object-based capabilities and some similar thing might be useful to let a VM make remote requests without giving it API keys. (Keys could be exfiltrated via prompt injection.)
- benswerd 6mo agoSo we have there are 3 solutions to this, Freestyle supports 2 of them: 1. Freestyle supports multiple linux users. All linux users on the VM are locked down, so its safe to have a part of the vm that has your secret keys/code that the other parts cannot access. 2. A custom proxy that routes the traffic with the keys outside 3. We're working on a secrets api to intercept traffic and inject keys based on specific domains and specific protocols starting with HTTP Headers, HTTP Git Authentication and Postgres. That'll land in a few weeks.
- TheTaytay 6mo agoWow, forking memory along with disk space this quickly is fascinating! That's something that I haven't seen from your competitors. If the machine can fork itself, it could allow for some really neat auto-forking workflows where you fuzz the UI testing of a website by forking at every decision point. I forget the name of the recent model that used only video as its latent space to control computers and cars, but they had an impressive demo where they fuzzed a bank interface by doing this, and it ended up with an impressive number of permutations of reachable UI states.
- benswerd 6mo agoThat’s what I’m hoping for!
- laalshaitaan 6mo agoI think one of the very few who actually support ebpf & xdp, which you do need when you're building low level stuff. + the bare metal setup is like out of the world lol.
- benswerd 6mo agoTx it took a lot of work lol
- jFriedensreich 6mo agoNon open source and non local SAAS sandboxes are offensive to even try to launch. No one needs this and the only customers will be vibe coders who just don't know any better. There are teams building actual sandboxes like smolmachines, podman, colima and mre. At least be honest and put the virtualisation tech you are using as well as that its closed source SAAS on the landing page to safe people time.
- benswerd 6mo agoOur users are platforms, and many of the best already build on us. Self hosting is a valuable feature but our technology is unfriendly to small nodes — it will not work on consumer hardware. Many of the optimizations we spend our time on only seriously kick in above 2TB of storage and above 500GB of RAM.
- senko 6mo ago> Non open source and non local SAAS sandboxes are offensive to even try to launch. No one needs this and the only customers will be vibe coders who just don't know any better. This is simply not true, but also not a very charitable take.
- alasano 6mo agoYour comment could have been "I prefer these open source alternatives:" but you chose to be a hater. There's nothing wrong with offering services that people find useful.
- brap 6mo agoApparently it’s offensive to even try to make things people want
- csomar 6mo ago> No one needs this VMWare was acquired for $69Bn.
- qainsights 6mo agoIs this similar to https://instavm.io/ https://instavm.io/?
- benswerd 6mo agoNever tried them, I think the weird thing about VM providers is the difference really all is in the execution. These guys seem great in concept but I don’t know enough about how they properly work.
- thepoet 6mo agoHi Ben, one of the founders of InstaVM here. Congrats on the launch! Here is a feature tour of InstaVM https://instavm.io/blog/meet-instavm-infra-for-your-agents https://instavm.io/blog/meet-instavm-infra-for-your-agents We would be publishing on the tech soon. Would love to give you a demo of InstaVM and trade notes. Let me know abhishek@instavm.io
- tuhgdetzhh 6mo ago[dead]
- jbethune 6mo agoCongratulations on the launch! Will definitely test this out.
- nyellin 6mo agoIs it possible to run a Kubernetes cluster inside one? (E.g. via KIND.) If so, we'd very much like to test this. We make extensive use of Claude Code web but it can't effectively test our product inside the sandbox without running a K8s cluster
- benswerd 6mo agoYes! You can def run something like K3s in these VMs.
- fawabc 6mo agohow does this differ from daytona or e2b?
- ianberdin 6mo agoAlso modal.com, I saw a few more as well.
- benswerd 6mo agoGenerally compared to those two more powerful. Freestyle VMs are full Debian machines, with support for sysd, docker in docker, multiple users, hardware virtualization etc. Daytona and E2B are both great "sandbox" providers but don't really feel like VMs/you can't run everything you can in an EC2. We also support the forking/snapshotting/long running jobs that they struggle to.
- umarcyber 6mo agoYour UI design is really nice.
- messh 6mo agoCheckout shellbox.dev, you can do pretty much the same, automating it all bia ssh
- holoduke 6mo agoThe problem with agents is that it is currently way too expensive. 100 times more expensive maybe. Another big issue is the lack of interactivity with an agent. Therefor for now agentic development is only viable from your own machine. And there isolation is less of an issue easier to manage.
- _pdp_ 6mo agoNice work. However, 50 concurrent VMs is not a lot. Similar limits exists on all cloud providers, except perhaps in AWS where the cost is prohibitive and it is slow. Earlier this year, we ended up rolling out own. It is nothing special. We keep X number of machines in a warm pool. Everything is backed by a cluster of firecracker vms. There is no boot time that we care about. Every new sandbox gets vm instantaneously as long as the pool is healthy.
- benswerd 6mo ago50 is not heavy, what is heavy is 1000 VMs that can be paused/brought back 50 in 1 second. Though generally ya, handrolling this stuff can work at the scale of 50 VMs, it becomes a lot harder once you hit hundreds/thousands.
- kjok 6mo agoThanks for sharing your approach! > It is nothing special. We keep X number of machines in a warm pool. I'd love to better understand the unit economics here. Specifically, whether cost is a meaningful factor. The reason I ask is that many startups we've seen focus heavily on optimizing their technology to reduce cold/boot startup times. As you pointed out, perceived latency can also be improved by maintaining a warm pool of VMs. Given that, I'm trying to determine whether it's more effective to invest in deeper technical optimizations, or to address the cold start problem by keeping a warm pool.
- ianberdin 6mo agoCongrats guys! Would share some technical details, I bet you have great stories to tell. Let’s, what is forking? You completely copy disk, make ram snapshot and run it? If CoW, but ram? You mentioned 8GB ram vms. Sounds like impossible to copy 8Gb under 500ms, also disk?
- benswerd 6mo agoSo fork time is actually O(1) with VM size, its 500ms even for 64gb + disk. We're using some pretty weird COW techniques to pull it off.
- ianberdin 6mo agoInsane. Does it possible to fork to another bare metal machine? Maybe multi region as fly io. If not, I bet you have huge disk sizes on your machines to store all the snapshots (you said, you store them and bill only for disk space).
- benswerd 6mo agoSo forking across multiple nodes in that speed is not possible — we run extremely beefy nodes in order to avoid moving VMs across nodes as much as possible. We are researching systems of hot moving VMs across VMs but it would have very different performance characteristics.
- ianberdin 6mo agoYeah, I see. Is it possible to get a corrupted state? Let’s say we had realtime database actively writing at that moment?
- benswerd 6mo agoIt is impossible. Our tech is not decades old so there is a chance we've missed something but our layer management is atomic so I'd be shocked if you'd be able to corrupt state across forks/snapshots.
- cheema33 6mo agoI currently use lightweight VMs (Proxmox containers) and git worktrees. I can fork an existing VM in in seconds. It is not entirely clear to me what I would gain from using your solution.
- benswerd 6mo agoProxmox forking in a few seconds is a miracle! These are likely only a better value for you at large scale/if you start wanting to run hundreds.
- nightrate_ai 6mo ago[dead]
- aaztehcy 6mo ago[flagged]
- maltyxxx 6mo ago[dead]
- lawrencechen 6mo agoCan you develop freestyle in freestyle vms?
- benswerd 6mo agoYessir, we haven't mastered it yet but we've compiled the kernel with enough flags for stuff like nftables and KVM to make it possible.
- alasano 6mo agoJust want to say that even if alternatives exist (not necessarily exact capabilities obviously), I appreciate what seems to be genuine excitement on your part of having built something cool / best in class. So best of luck with your vision for it!
- atlasagentsuite 6mo ago[flagged]
- sn9 6mo agoWhat are some examples of this?
- benswerd 6mo agoCI Builders/QA Agents can do this very well. User session starts, bring VM up with content + dependencies, when session is done throw it away. Keeps it clean, debuggable, fast and cheap.
- benswerd 6mo agoFreestyle has really built with this in mind. We propose a primary architecture built around declarative configuration of the vm with a git repo as external source of truth. If the VM crashes/you have another idea/you want to try something else it should be reconstructable from outside of the VM. However, I think this is potentially unrealistic. While it is the ideal architecture, I hear more and more every day people who just want to have the VMs run for months at a time.
- siscia 6mo agoIt is not clear to me how much CPU I get. "Unlimited" as in 8vCPU and then I am billed for it on consumption?
- benswerd 6mo agoBilled for wall time. whichever plan you are on you get in credits, so hobby plan gets $50 of credits and beyond that billed on per CPU wall time.
- Bnjoroge 6mo agoLooks cool - would be great to see a PR with some benchmarks on this repo if you can: https://github.com/computesdk/benchmarks https://github.com/computesdk/benchmarks edit: just saw the pr for freestyle. something seems to be blocking, but curious how it compares: https://github.com/computesdk/benchmarks/pull/41 https://github.com/computesdk/benchmarks/pull/41
- Shin0221 6mo ago[flagged]
- deleted 6mo ago[deleted]
- esseph 6mo ago> In order to make this possible, we’ve moved to our own bare metal racks. Early in our testing we realized that moving VMs across cloud nodes would not have acceptable performance properties. We asked Google Cloud and AWS for a quote on their bare metal nodes and found that the monthly cost was equivalent to the total cost of the hardware so we did that. Yes! And good on you, well-tuned bare metal performance is hard to beat.
- k38f 6mo ago500ms fork of a running VM with full memory state is the kind of thing I'd assume wasn't possible until I saw it work. What does failure look like — does the fork just not happen, or can you get partial state?
- benswerd 6mo agoThere is no partial state really possible. We can run out of space on a Node and just say no. But the nature of memory forking is if you don't literally do it 100% right it crashes immediately (I know cuz it took me a while too get it right).
- psychomfa_tiger 6mo ago[flagged]
- benswerd 6mo agoTBH I wouldn't recommend using it for this. I'm a big believer in agent chat running outside of the VM, where you can get much better control over the chat loop. I would treat the VM as a tool the agent is using rather than the agent's environment. Like the agent is a human using a machine and watching it, rather than trying to watch it from inside the machine. Then there are great existing observability tools, my fav is langfuse.
- brap 6mo agoBut doesn’t this defeat the purpose? I would actually imagine this would be useful for observably in the sense that you can fork and then kill the loop in the fork, hop into an interactive session to figure out what it’s doing, while the loop is still running in the original instance.
- benswerd 6mo agoI don't believe so. while it is technically easy to fork claude code running in these VMs, its not technically difficult to fork a conversation loop outside of the VM as well. What matters is that its all forked atomically, which can be done with resources outside of the VM as well.
- brap 6mo agoFair enough, and I respect you pointing out the alternatives
- lukebaze 6mo agoThe observability point is real but honestly the loop detection problem is more about how you structure your agent than the sandbox. When I've had agents go rogue, the issue was always the outer loop logic, not visibility into the VM. What does your current loop controller look like?
- MeetRickAI 6mo ago[dead]
- brap 6mo agoVery nice, congrats! One thing: >Freestyle is the only sandbox provider with built-in multi-tenant git hosting — create thousands of repos via API and pair them directly with sandboxes for seamless code management. Maybe I’m just stupid, but I don’t know what this means. I initially thought I’m your target audience but after failing to understand this part I’m thinking maybe I’m not? I honestly don’t know.
- benswerd 6mo agoIf git isn't for you we'd still love to support you. We believe to build the sandboxes for coding agents you also need to provide git repos for them so we do that as well. You can easily say give me this vm with these 3 repos and these permissions with us. But that said, the sandbox stands on its own without it.
- ewidar 6mo agoI think you should just explain that part more clearly: why would they want you to host git repos on their behalf.
- benswerd 6mo agoFreestyle isn't designed for an individual engineer working on their Github repos. Its designed for platforms building coding agents that want to take the place of Github all together. Those platforms need some source of truth alongside the VMs, just like how you don't store all of your important documents on your personal computer. That is why we offer git.
- stingraycharles 6mo agoIt’s difficult to understand the content and what the product actually is, as I’ve also mentioned in another reply. I think the product is probably great, but you need to improve the communication, it’s too abstract. I don’t know what “give each sandbox a unique git repository” does for me in practice, what problem it solves. You’re not providing any practical problems your product is intended to solve.
- zhdhdjfdhsbs 6mo agoWqq wwiq and hdhddjdbnzzs S
- areys 6mo ago[flagged]
- bhaktatejas922 6mo agodo you think the industry is overfixated on startup times? what are better metrics people building with sandboxes should pay attention to
- benswerd 6mo agoSo first I don't, I think startup times are fundamentally really important. 5s is different than 1s is different than 500ms is different than 200ms and users notice. I don't think people run real world benchmarks on what that coldstart really means though, like time to first response from a NextJS is a very important benchmark for Freestyle and we've spent a lot of time on it. While Daytona sandboxes boot faster than Freestyle ones our first response is an order of magnitude ahead of theirs. I think another important one is concurrency: In worst case scenarios how many VMs can you get from a provider in a 5 second period is important. I also think not enough time is spent on "Does it actually work on this VM", stuff like postgres, redis, ntftables, complex linux binaries that are hard to run need to work on these sandboxes because AI is going to need them and I don't think there has really been a feature-bench system yet. Networking/snapshotting/persistence characteristics all also need to come into this. I
- csomar 6mo agoI was intrigued to try but your web app is so extremely slow, it takes up to 30+ seconds to move from one tab to the next. Not exactly selling your point of being a super fast provisioning service. Another thing I am wondering. You seem to be selling this as VMs configurable from node/bun. Wouldn't a CLI make more sense here? Another question: How hard do you think it'll be to integrate this with something like Claude Code. ie: /resume in claude code both return your session and wake up your vm. Or even better /resume from freestyle and have your claude code session open where you left it.
- benswerd 6mo agoI'm not sure what you saw as slow, I'd love to improve it. Do you mean the dashboard? We're built as an API for platforms to build on rather than tool for individual developers. Oriented at platform orchestrating tens of thousands at VMs rather than individuals using CLI. We also have a CLI but its primarily a debugging and testing tool. Resuming a freestyle VM with claude code in it will just work. You can do that via SSH.
- csomar 6mo ago> I'm not sure what you saw as slow, I'd love to improve it. Do you mean the dashboard? Switching tabs in the Dashboard (Domains/Routes/etc.) was basically unusable about 4 hours ago. It's noticeably better now, though there's still some latency (just retested).
- randomtoast 6mo agoDo you have any recommendations for CLI-based microVM solutions that support running multiple instances of Claude Code with "--yolo sandboxing" on Linux?
- techpulselab 6mo ago[dead]
- CompuIves 6mo agoThis is really cool to see, reminds me of the early days of CodeSandbox. Though this API looks _fantastic_. I love that you do VM configuration using `with`.
- benswerd 6mo agoWe read your blogs when building all of this!
- AlexanjaSenke 6mo ago[dead]
- sonink 6mo agoCongratulations on the launch ! We run upwards of a thousand sandboxes for coding agents - but these are all standard VM's that we buy off the shelf from Azure, GCP, Akamai and AWS. I am not sure why we should use this instead of the standard VM's? Pricing could be one part, but not sure if the other features resonate. Forking is interesting, but I would need to know how it works and if it is in the blast radius of the agent execution. If we need to modify the agent to be cognizant of forking, then that is a complexity which could be very expensive to handle in terms of context. If not, then I am not sure what is the use for it. Sandbox start time at 500ms is definitely interesting. But its something we already are on track to reproduce with a pooled batch of VM's. So not sure if that in itself is worth paying for the premium. My two cents on the space is that agents are rapidly becoming more capable to just use the tooling developed for humans. All clouds provide a CLI which agents can already use to orchestrate - they should just use the VM's designed for humans through the CLI. Our agent can already 'login' to any VM on the cloud and use the shell exactly like a human would. No software harness is required for this capability. The agent working on a VM is indistinguishable from humans.
- bac2176 6mo ago[flagged]
- mt18 6mo ago[dead]
- orliesaurus 6mo agoThere are many providers popping up every day offering sandboxes, I think Cloudflare is ahead of the game for pricing and performance, that being said it would be super nice to see a huge competitor analysis: Cloudflare vs e2b vs daytona vs freestyle vs whatever else
- philbitt 6mo ago[dead]
- Sattyamjjain 6mo ago[flagged]
- etse 6mo agoThe memory forking is really interesting. I wonder if copy-on-write at the VM level, O(1) with respect to machine size, won't scale cost with how many forks to take, but 320ms median seems good for the branch-and-explore pattern without reprovisioning every time. One gap I'm noticing in these comments and in the current sandbox landscape is Windows. Every platform mentioned in these comments like E2B, Daytona, Fly Sprites, Sandflare appears Linux-native. Makes sense for coding agents targeting Debian environments, but a real category exists to automate Windows-specific workflows: enterprise software, ERP systems, anything that runs only on Windows. If anyone wants to run agents in Mac or Linux and need to access Windows for computer use, Dexbox could be helpful. [github.com/getdexbox/dexbox] I launched an open source developer tool called Dexbox to run agent workloads that quickly provision and run Windows desktops. It's a CLI and MCP experience that's different from Freestyle, but slightly closer to our Windows-specific production infra, Nen. I like Freestyle's cool UI that shows off the unique technical approach and developer friendliness. Nen's a bit closer to that experience.
- BlueRock-Jake 6mo agoTon of people have mentioned this but what you're doing with memory forking is pretty unique. Most sandboxes seem to just fork the filesytem and call it a day. Forking full VM memory mid-exec is taking it to another level entirely. Would be very interested to hear how the implementation looks under the hood, specifically how you handle dirty memory pages across forks without the pause ballooning.
- studio-m-dev 6mo ago[flagged]
- danielhanchen 6mo agoCongrats on the launch!
- willamhou 6mo ago[dead]