4 ms·
Running agents in a sandbox or VM is the wrong pattern
- zrail 26d ago> A worker reads the tail of the log, performs exactly one step—either inference or tool execution—appends the result, and forgets everything. It holds zero session state between turns. My current understanding of an agent session is that the entire session (modulo compaction, reordering tools, etc) is sent to the stateless inference engine every time inference happens. The quoted section doesn't seem workable unless "read the tail of the log" actually means "read the entire session log". There are ways to make that work but workers being completely stateless makes it hard.
- xij 26d agoFirst, the worker has to read the entire session log only if the next action is inference. You are right that my writing is misleading. To clarify that, I mean the worker only needs the tail of the log to decide what to do next: inference or execute tools.
- zrail 25d agoThat does make more sense, thanks for the clarification.
- pascalfenkam 25d agoI have been thinking alot about how to use agent to operate infrastructure. When I saw the title I was curious as to how would you operate agents if not in a sandbox or VM. Reading the actual article I understand that you are referring to personal agents. Is that correct?
- xij 25d agoQuite the opposite, the personal agents space is the one that we are trying to avoid. There are many companies working on personal agents, or in a broader term, we can call them the interactive agents. However, just like the internet, not all computational work is interactive. Batch processing is also very important. This is the motivation for me to write this blog and build this project. For example, one of our customers builds a swarm of agents that simulate customer behavior to test the products before launch. Such a task can be run in batch and doesn't require a human to sit with the agents. (Therefore, non-interactive)
- coder-pm 23d agoThis makes sense for stateless workers, which don’t have to keep the context between the steps. What about the interactive agents, holding ssh session or repository state between the steps? That’s a different case, isn’t it?