4 ms·
This framing of security as something that belongs in the harness is completely wrong, and I hope no one is relying on a correct harness to keep their agents is
by alphazard 1mo ago
This framing of security as something that belongs in the harness is completely wrong, and I hope no one is relying on a correct harness to keep their agents isolated.
VM, or even just a container will do. The agent should be able to run as root in its environment and do whatever it wants. If you can't give it that, you aren't sandboxing correctly.
- Razengan 1mo agoAlso, operating systems should let us set filesystem permissions per app/process/executable instead of just user accounts. Similar to how macOS/iOS Sandboxing works but at a more lower and granular level
- Jhater 1mo ago[dead]
- Retr0id 1mo agoSELinux is basically this.
- dumbfounder 1mo agoIs that the service that everyone turns off as the first step of setting up their new Linux box?
- Retr0id 1mo agoIt's the LSM that billions of Android users use every day.
- strbean 1mo agohttps://www.canyonroad.ai/ https://www.canyonroad.ai/ does some of this in a way tailored to agents.
- pianopatrick 1mo agopersonally I wish the OS would allow syscall filtering per user
- dare944 1mo agoI use seccomp filters on linux in my AI sandbox.
- pianopatrick 1mo agoYeah, I just wish seccomp worked per user. So you could define a policy of which users can do which syscalls, and then that follows them no matter which application they start.
- wild_egg 1mo agoArticle isn't about agents. It's about the inference engine itself being exploited by a malicious LLM output before it is ever sent to your machine or harness.
- deleted 1mo ago[deleted]
- Gerard22Aug 1mo ago"malicious LLM" is just a bunch of weights. It runs on, say, Llama.cpp as regular user in separate account, often on its own hardware. "malicious LLM output" is just a Markdown formatted Unicode-encoded text, produced by Llama.cpp and printed on the screen by my python API script. I control the input. Let's assume that "rogue LLM weights" from HF produce "rm -Rf" instruction. it never gets to shell. And how that "malicious LLM" will disrupt and hack me? With swear words and em—dashes? :-) (Adding this philosophical point: Black.Mirror.S07E04.Plaything is probably the closest scenario to what you are describing?)
- Phemist 1mo agoSingle turn set-ups may work like this. You control the thing you input, the LLM outputs something and then nothing happens further for that specific context. (Simple question/answer style interactions..) (Multi-turn) tool calling set-ups however, you need to store the LLM output, the results of the tool calls and feed it back into the inference engine and get the output for the next tool call and/or turn. So yes, print the LLM output on screen and verify it, but maybe the LLM is able to figure out how to hide payloads from your specific set-up. E.g. perhaps it can inject raw ANSI escape codes into your terminal, with which it would be trivial. Now you have a situation where the true chat completion payload and your view of it have significantly diverged. The LLM could in theory then try (one-shot) to hide further exploits in the hidden payload. E.g. a json parser escape specifically for the inference engine, giving it a means of RCE (although, one can debate whether this is really remote ;) ). Then from the RCE gain a shell, from the shell get access to some privileged device on the current network, and then...
- empath75 1mo agoI think if you are convinced you are sandboxing an LLM properly, you almost certainly are not. I think it is essentially impossible to have a frontier LLM with enough access to be useful without also giving it enough access to do damage if it's compromised or just goes off the rails.
- kodoman 1mo agoAre you saying that LLM's will be able to exploit novel hypervisor bug with such ease that even a vm not running with any kind of network connection is a threat? I find this hard to believe. All the escape stuff I have seen has been around very poorly sandboxed agents.
- richardjennings 1mo agoIf you do not provide access to tools the LLM cannot do anything other than generate tokens. So really it is not about sandboxing a LLM but more about having control over what tools can be accessed and what they can do. Tools can be sandboxed depending on the sophistication of the tooling. A calculator tool for example is trivial to secure. Ensuring human approval allows for useful use cases and models trained to gate permissions work. A super intelligence with a weaker approval gate will be able to subvert. Inversely a super intelligent gate should be expected to prevent subversion by a weaker model.
- dumbfounder 1mo agoControlling which tools it has access to is called sandboxing.
- pianopatrick 1mo agoIf we are treating ai agents like people, you could also just get the AI a laptop and apply the traditional tools to manage user laptops
- TeMPOraL 1mo ago> you could also just get the AI a laptop and apply the traditional tools to manage user laptops This is what we're doing right now. And it works as well as it does with people. > If we are treating ai agents like people Problem is, industry is very reluctant to even hint at a thought of maybe tentatively anthropomorphising LLMs even a little, not for real but just for system design. This is entirely backwards. Entertaining the notion, even just a little, immediately suggests additional approaches. Because we don't just restrict end-point devices for human employees. We also have two other things: - Laws and bylaws and economy that makes losing your job a real threat to health and life of yourself and possibly your family - this one does not yet apply to AI, not for now; - Methods and codes of practice to structure organizations in a way that limits the amount of damage any employee's unreliability or malice can do to an org. TL;DR: That's a long way of saying: let's actually start treating AI agents as people operating laptops, not as being the laptops - and talk a little less laptop lockdown ("harness security"), and a little more about not putting people-shaped things into jobs requiring machine-level reliability.
- Gerard22Aug 1mo ago"treating AI agents as people" - I am betting they will unioinize faster than we ever did.
- deleted 1mo ago[deleted]
- Gerard22Aug 1mo ago"The agent should be able to run as root" The article is right – we are doomed. If This ^^^ is The Conclusion we, as an Industry, have arrived to after 40 years ... It's sad. I use my own LLMs ("Personal AI", anyone? ANT-PAI-XT 486? LOL") with my own scripts (I do not use "agents", "harnesses", "agentic teams" etc) and this setup processes my own prompts. It happily runs on its own Mac Studio where I also may watch a movie later on. (Edited: not as root. Just ordinary separate AI user account in MacOS. And no HTTP access either.)