3 ms·
> Getting the most out of models requires maximally exploiting parallelism, so sandboxing is a must. What are your thoughts on checkpointing as a refinement of
by jsnell 1y ago
> Getting the most out of models requires maximally exploiting parallelism, so sandboxing is a must.
What are your thoughts on checkpointing as a refinement of sandboxing? For tight human/llm loops I find automatic checkpoints (of both model context and file system state) and easy rolling back to any checkpoint to the most important tool. It's just so much faster to roll back on major mistakes and try again with the proper context, than to try to get the LLM to fix a mistake, since now the broken code and invalid assumptions are contaminating the context.
But that relies on the human in the loop deciding when to undo. Are you giving some layer of your system the power of resetting sub-agents to previous checkpoints, do you do a full mind-wipe of the sub-agents if they get stuck and try again, or is the context rot just not a problem in practice?
- mike_hearn 1y agoI want to minimize human in the loop. At the moment my agent allows user interaction in one place, after a spec is written+reviewed+updated to reflect the review, it stops. You can then edit the spec before asking for an implementation. It helps to catch cases where the instructions were ambiguous. At the moment my agent is pretty basic. It doesn't detect endless loops. The model is allowed to bail at any time when it feels it's done or is stuck, so it doesn't seem to need to. It doesn't checkpoint currently. If it does the wrong thing you just roll it all back and improve the AGENTS.md or the mission text. That way you're less likely to encounter problems next time. The downside is that it's an expensive way to do things but for various reasons that's not a concern for this agent. One of the things I'm experimenting with is how very large token budgets affect agent design.