3 ms·
Ask HN: How do you prevent AI agents from going rogue in production?
Hi all!
There seems to be an ongoing trend (and my gut feeling) of companies moving from chatbots to AI agents that can actually execute actions—calling APIs, modifying databases, making purchases, etc.
I'm curious: if you're running these in production, how are you handling the security layer beyond prompt injection defenses?
Questions:
- What stops your agent from executing unintended actions (deleting records, unauthorized transactions)?
- Have you actually encountered a situation where an agent went rogue, and you lost money or data?
- Are current tools (IAM policies, approval workflows, monitoring) enough, or is there a gap?
Trying to figure out if this is a real problem worth solving or if existing approaches are working fine.
- Agent_Builder 9mo ago[dead]
- techbuilder4242 9mo agoThis is a great insight, thank you for sharing! A few follow-ups if you don't mind: - When you say "tightening execution boundaries," are you doing this at the orchestration layer (LangChain/CrewAI/custom), or did you build middleware that sits between the agent and APIs? - How do you handle the tradeoff between narrow permissions per step vs. agent flexibility? - For "step-level control and visibility gap" - that is the most impactful insight. I'm trying to wrap my head arond this particular one. Looks like that sooner or later this gap may be addressed by current AI models providers.