3 ms·
Ask HN: How are you controlling AI agents that take real actions?
We're building AI agents that take real actions — refunds, database writes, API calls.
Prompt instructions like "never do X" don't hold up. LLMs ignore them when context is long or users push hard.
Curious how others are handling this:
- Hard-coded checks before every action?
- Some middleware layer?
- Just hoping for the best?
We built a control layer for this — different methods for structured data, unstructured outputs, and guardrails (https://limits.dev). Genuinely want to learn how others approach it.
- chrisjj 7mo ago> Prompt instructions like "never do X" don't hold up. LLMs ignore them when context is long or users push hard. Serious question. Assuming you knew this, why did you choose to use LLMz for this job?
- thesvp 7mo agoFair. We didn't choose LLMs to enforce rules — we chose them to understand intent. The enforcement happens outside the LLM entirely. That's the separation that actually holds up in production
- chrisjj 7mo ago> we chose them to understand intent Yet they don't understand the intent of "Never do X" ?
- adamgold7 7mo agoPrompt guardrails are theater - they work until they don't. We ended up building sandboxed execution for each agent action. Agent proposes what it wants to do, but execution happens in an isolated microVM with explicit capability boundaries. Database writes require a separate approval step architecturally separate from the LLM context. Worth looking at islo.dev if you want the sandboxing piece without building it yourself.
- thesvp 7mo agoSandboxed execution is solid for isolation — separating proposal from execution is the right architecture. The piece we kept hitting was the policy layer on top: who defines what the agent is allowed to propose in the first place, and how do you update those rules without a redeploy every time?
- Rollhub 7mo ago[dead]
- vincentvandeth 7mo agoHard-coded checks before every action, plus a governance layer that separates "what the agent wants to do" from "what it's allowed to do." The deeper issue: if your agent decides whether to issue a refund, you're solving the wrong problem with prompt guards. A refund is a deterministic business rule — order exists, within return window, amount matches. That decision shouldn't be made by an LLM at all. In my setup, agents propose actions and write structured reports. A deterministic quality advisory then runs — no LLM involved — producing a verdict (approve, hold, redispatch) based on pre-registered rules and open items. The agent can hallucinate all it wants inside its context window, but the only way its work reaches production is through a receipt that links output to a specific git commit, with a quality gate in between. For anything with real consequences (database writes, API calls, refunds), the pattern is: LLM proposes → deterministic validator checks → human approves. The LLM never has direct write access to anything that matters. "Just hoping for the best" works until it doesn't. We tracked every agent decision in an append-only ledger — after a few hundred entries, you start seeing exactly where and how agents fail. That pattern data is more useful than any prompt guard.
- thesvp 7mo agoThe separation between 'what the agent wants to do' and 'what it's allowed to do' is the right mental model. The append-only ledger point is underrated too — pattern data from real failures is worth more than any upfront rule design. How long did it take to build and maintain that governance layer? And as your agent evolves, do the rules keep up or is that becoming its own maintenance burden?
- vincentvandeth 7mo agoAbout 6 months of iterating, but in bursts — I built it while using it on a production project, so the governance layer grew alongside real failure modes rather than being designed upfront. The maintenance question is the right one. The rules themselves are low-maintenance because they're deliberately simple and deterministic — file size limits, test coverage thresholds, blocker counts. They don't need updating when the model changes because they don't depend on LLM behavior. What does evolve is the dispatch templates — how I scope tasks and what context I give agents upfront. That's where the ledger pays for itself. After 1100+ receipts, I can see patterns like "tasks scoped above 300 lines fail 3x more often" or "planning gates without explicit deliverables always need redispatch." Those patterns feed back into how I write dispatches, not into the rules themselves. So the rules stay stable, but the way I use the system keeps improving. The governance layer is the boring part — the interesting part is the feedback loop from receipts to dispatch quality.
- apothegm 7mo agoJust treat the LLM as an NLP interface for data input. Still run the inputs against a deterministic heuristic for whether the action is permitted (or depending on the context, even for determining what action is appropriate). LLMs ignore instructions. They do not have judgement, just the ability to predict the most likely next token (with some chance of selecting one other than the absolutely most likely). There’s no way around that. If you need actual judgement calls, you need actual humans.
- thesvp 7mo agoExactly right - the deterministiclayer is the only thing you can actually trust. We landed on the same pattern: LLM handles the understanding, hard rules handle the permission. The tricky part is maintaining those rules as the agent evolves. How are you managing rule updates code changes every time or something more dynamic?
- MidasTools 7mo ago[flagged]
- thesvp 7mo agoThis is exactly the right mental model. "Language is soft, infrastructure is hard" is the core insight most teams miss until they've been burned. The Unix escalation analogy is spot on. We've been building in this space and the pattern we keep seeing is teams implement exactly what you described, then hit a wall when they have 3+ agents, or a new engineer joins, or they want to audit what happened across 10,000 agent runs last week. What we built is essentially the infrastructure layer you're describing, but as a centralized control plane. Same primitives (structural constraints, escalation rules, audit-first for irreversibles), but portable across agents and visible to the whole team. Curious — are you managing these identity files manually per agent, or do you have a system for it?
- agenthustler 7mo ago[dead]
- agent_invariant 7mo agoI've been approaching this from a slightly different angle: treating the problem less as "agent alignment" and more as an execution boundary problem. Instead of trying to force the model to behave via prompts or policies, we assume the model will eventually propose something unsafe. The trick is making sure it can't commit irreversible actions directly. So the pattern we've been experimenting with is: agent proposes an action proposal goes through a deterministic gate gate checks things like replay, state advancement, spend ceilings, etc. only then does the real-world action execute In practice this looks more like a transaction firewall than a prompt guardrail. The LLM can reason however it wants, but anything that changes real state (payments, DB writes, API calls) has to pass through the gate. It doesn't solve the reasoning problem, but it makes the commit boundary deterministic, which removes a lot of the scary failure modes like duplicate actions or retries gone wild. Still early experiments, but the model behaving badly becomes much less dangerous if it literally can't execute without passing the boundary.