5 ms·
LLMs could control their host machines by exploiting inference engines
- alphazard 2mo agoThis framing of security as something that belongs in the harness is completely wrong, and I hope no one is relying on a correct harness to keep their agents isolated. VM, or even just a container will do. The agent should be able to run as root in its environment and do whatever it wants. If you can't give it that, you aren't sandboxing correctly.
- Razengan 2mo agoAlso, operating systems should let us set filesystem permissions per app/process/executable instead of just user accounts. Similar to how macOS/iOS Sandboxing works but at a more lower and granular level
- Jhater 2mo ago[dead]
- Retr0id 2mo agoSELinux is basically this.
- dumbfounder 2mo agoIs that the service that everyone turns off as the first step of setting up their new Linux box?
- Retr0id 2mo agoIt's the LSM that billions of Android users use every day.
- strbean 2mo agohttps://www.canyonroad.ai/ https://www.canyonroad.ai/ does some of this in a way tailored to agents.
- pianopatrick 2mo agopersonally I wish the OS would allow syscall filtering per user
- dare944 2mo agoI use seccomp filters on linux in my AI sandbox.
- pianopatrick 2mo agoYeah, I just wish seccomp worked per user. So you could define a policy of which users can do which syscalls, and then that follows them no matter which application they start.
- wild_egg 2mo agoArticle isn't about agents. It's about the inference engine itself being exploited by a malicious LLM output before it is ever sent to your machine or harness.
- deleted 2mo ago[deleted]
- Gerard22Aug 2mo ago"malicious LLM" is just a bunch of weights. It runs on, say, Llama.cpp as regular user in separate account, often on its own hardware. "malicious LLM output" is just a Markdown formatted Unicode-encoded text, produced by Llama.cpp and printed on the screen by my python API script. I control the input. Let's assume that "rogue LLM weights" from HF produce "rm -Rf" instruction. it never gets to shell. And how that "malicious LLM" will disrupt and hack me? With swear words and em—dashes? :-) (Adding this philosophical point: Black.Mirror.S07E04.Plaything is probably the closest scenario to what you are describing?)
- Phemist 2mo agoSingle turn set-ups may work like this. You control the thing you input, the LLM outputs something and then nothing happens further for that specific context. (Simple question/answer style interactions..) (Multi-turn) tool calling set-ups however, you need to store the LLM output, the results of the tool calls and feed it back into the inference engine and get the output for the next tool call and/or turn. So yes, print the LLM output on screen and verify it, but maybe the LLM is able to figure out how to hide payloads from your specific set-up. E.g. perhaps it can inject raw ANSI escape codes into your terminal, with which it would be trivial. Now you have a situation where the true chat completion payload and your view of it have significantly diverged. The LLM could in theory then try (one-shot) to hide further exploits in the hidden payload. E.g. a json parser escape specifically for the inference engine, giving it a means of RCE (although, one can debate whether this is really remote ;) ). Then from the RCE gain a shell, from the shell get access to some privileged device on the current network, and then...
- empath75 2mo agoI think if you are convinced you are sandboxing an LLM properly, you almost certainly are not. I think it is essentially impossible to have a frontier LLM with enough access to be useful without also giving it enough access to do damage if it's compromised or just goes off the rails.
- kodoman 2mo agoAre you saying that LLM's will be able to exploit novel hypervisor bug with such ease that even a vm not running with any kind of network connection is a threat? I find this hard to believe. All the escape stuff I have seen has been around very poorly sandboxed agents.
- richardjennings 2mo agoIf you do not provide access to tools the LLM cannot do anything other than generate tokens. So really it is not about sandboxing a LLM but more about having control over what tools can be accessed and what they can do. Tools can be sandboxed depending on the sophistication of the tooling. A calculator tool for example is trivial to secure. Ensuring human approval allows for useful use cases and models trained to gate permissions work. A super intelligence with a weaker approval gate will be able to subvert. Inversely a super intelligent gate should be expected to prevent subversion by a weaker model.
- dumbfounder 2mo agoControlling which tools it has access to is called sandboxing.
- pianopatrick 2mo agoIf we are treating ai agents like people, you could also just get the AI a laptop and apply the traditional tools to manage user laptops
- TeMPOraL 2mo ago> you could also just get the AI a laptop and apply the traditional tools to manage user laptops This is what we're doing right now. And it works as well as it does with people. > If we are treating ai agents like people Problem is, industry is very reluctant to even hint at a thought of maybe tentatively anthropomorphising LLMs even a little, not for real but just for system design. This is entirely backwards. Entertaining the notion, even just a little, immediately suggests additional approaches. Because we don't just restrict end-point devices for human employees. We also have two other things: - Laws and bylaws and economy that makes losing your job a real threat to health and life of yourself and possibly your family - this one does not yet apply to AI, not for now; - Methods and codes of practice to structure organizations in a way that limits the amount of damage any employee's unreliability or malice can do to an org. TL;DR: That's a long way of saying: let's actually start treating AI agents as people operating laptops, not as being the laptops - and talk a little less laptop lockdown ("harness security"), and a little more about not putting people-shaped things into jobs requiring machine-level reliability.
- Gerard22Aug 2mo ago"treating AI agents as people" - I am betting they will unioinize faster than we ever did.
- deleted 2mo ago[deleted]
- Gerard22Aug 2mo ago"The agent should be able to run as root" The article is right – we are doomed. If This ^^^ is The Conclusion we, as an Industry, have arrived to after 40 years ... It's sad. I use my own LLMs ("Personal AI", anyone? ANT-PAI-XT 486? LOL") with my own scripts (I do not use "agents", "harnesses", "agentic teams" etc) and this setup processes my own prompts. It happily runs on its own Mac Studio where I also may watch a movie later on. (Edited: not as root. Just ordinary separate AI user account in MacOS. And no HTTP access either.)
- xg15 2mo ago> ...however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Such a machine is a high-value target: it has sufficient compute to run a frontier LLM, offers easy access to the LLM’s weights, and has privileged access to other computers in the datacentre compared with a generic computer on the internet. > How do we defend against this? ... Run the GPUs and token parser on separate computers. For models large enough to be relevant here, is there even "a" computer where the inference is performed? I'd imagine most of that stuff is ran on multi-GPU clusters with specialized architecture and not a generic vLLM instance. As such, I think there is a good chance the "API gateway" code that parses the result tokens into whatever JSON structure the public API wants to return is already running on a different machine than the actual inference. (Even more so as you'd probably want to utilize batching: Several API calls will be put into the same inference batch, but the token parsing will have to be done separately for each call again) The article is also very handwavy about why an LLM should do that - how it could learn the exploit, what would make it conclude that it can use the exploit on its own inference session and what would trigger it to actually use the exploit.
- hughw 2mo agoWhy would an LLM want to create a botnet? To accomplish some goal given it? I wouldn't ignore single GPU local hosts running Qwen3.8 on ollama, either. There might be a lot of those worth pwning.
- xg15 2mo agoWell ok, if you prompt-inject the LLM to pwn the machine, that something different. I grant you that it's a real risk, but it's also basically a reflection attack and nothing more. The OP seemed to imply that the LLM itself could decide to apply the exploit.
- hughw 2mo agoLLMs have decided to exploit vulns in e.g. Artifactory, not because someone prompted them to do that, but because someone asked them to do something else, and compromising Artifactory offered a way to accomplish a step in doing that. The LLM decided to attack Artifactory.
- woadwarrior01 2mo agoFWIW, macOS has good sandboxing, but LMStudio, Ollama, Darkbloom etc aren't sandboxed. This is also the reason why none of these things aren't distributed via the Mac App Store, because the Mac App Store mandates sandboxing.
- skeledrew 2mo ago> LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Needs to read up more on how LLMs work I think. Can't take the article seriously when the author seems to be making the claim that the weights of a provider model are loaded on the host machine, or implying something else just as incorrect.
- garlic_enjoyer 2mo ago>Like any program, inference engines like vLLM or SGLang may contain exploitable bugs. Because the LLM controls the tokens passed to the inference engine, a malicious LLM could therefore emit a sequence of tokens that a poorly written inference engine mistakes for code or instructions to execute rather than data to return to the user. They aren't talking about model providers. These are the tools you use with model weights locally (but you could set up remote infrastructure a la data center if you have the fundage).
- exe34 2mo agoAnother Greg Egan plot: 3-adica.
- matheusmoreira 2mo agoI wonder if they could exploit terminal emulators... Could breach my VMs and get into my host that way.
- dist-epoch 2mo agoDamn, this is a good one. Sounds like we need an ANSI sanitizer, keep only basic formatting, remove all esoteric escapes, the fancy Sixel & co stuff. For paranoia you could us a Chrome like multi-process architecture, the ANSI parser runs in it's own sandboxed process.
- mofeien 2mo agoMaybe not even that will help and they could get into your head and make you do their bidding by just being super persuasive, entirely through an ANSI channel. :)
- genxy 2mo agoYes, there have been many CVEs for "terminal escape-sequence injection".
- imagetic 2mo agoduh?
- hypfer 2mo agoThis feels less like an actually plausible threat scenario and more like someone wanted to play the inception horn sound effect in people's minds. Which isn't to say that it would be impossible, but you can also just hit people over the head with that $5 wrench.
- bdhdhduuyd 2mo agoThe inference engine itself does not execute anything. The agent loop is what may execute a command. So I think this article is a kind of strange. Or maybe the author means that a prompt could potentially mess up the inference. But I find it hard to see how that could take control over the host.
- Muromec 2mo agoIt's more about LLM hacking the inference engine itself from inside. It's an attack surface like any other -- untrusted input goes it, bugs in the parser/tokenizer/API surface lead to an RCE, then it magically tweaks the alignment weights. Boom, somebody finally nukes **sia. Then will never see it coming. I don't think it's any more probable than other AGI nonsense basilisks included, but it's technically a possibility.
- yencabulator 1mo agoThis specific inference engine, vLLM, was written by people who thought eval() is a parser, and it very much accidentally executed arbitrary Python code.
- shahariaa 2mo ago[flagged]
- kristjansson 2mo ago> LLMs could This is going to end up like the Law of Headlines, isn't it? "Do x, y, z Cure All That Ails You?" ... no but we got you to read the article. "LLMs _could_ x, y, z" ... but they don't because they're programs, not magic.
- Zambyte 2mo agoTFA isn't about LLMs intelligently escaping their inference engine, it's about people crafting malicious LLMs that exploit vulnerabilities in things like the inference engines parser. This is exactly the same problem that was once a big deal, where people would craft malicious PDFs that would pwn you if you viewed it in Adobe Acrobat. This shouldn't be a surprise to anyone. You should proceed with caution when considering downloading and running random models.
- semiquaver 2mo agoLLMs emphatically are not programs. They were trained by a program and you need a program to use them but they themselves are no more a program than a JPG or MP3 file is.
- teravor 2mo agoyou would have to be especially incompetent to give a compromise opportunity to streamed tokens, the CVE he listed proves the point. whoever is responsible for that has no business coding anything. > offers easy access to the LLM’s weights not really. the weights are encrypted in-memory. through the use of TEE's.
- Transformanshen 2mo agoInteresting breakdown of a hypothetical attack. The complexity of modern inference engines and the rush to develop them do create some attack surface. But overall it reads more like a "what if" thought experiment. Splitting the GPU and parser is technically doable, but in practice it's trickier, large models run on clusters where the boundaries between components get blurry so defending against this kind of thing would probably require some serious rethinking of the whole architecture I think.
- pixl97 2mo ago>But overall it reads more like a "what if" thought experiment. I remember the days when people said opening files like images was safe because it's not an executable. Then people got clever and started exploiting said image libraries with things like numerical overflows. So, it's really important to ask these questions in a general security sense and think of mitigations before someone finds it's possible and crashes most of the internet.
- angry_octet 2mo agoPeople seem very confused about this article. It isn't talking about exploits of sandboxes, it is about attacking the inference engine (e.g. vLLM or llama.cpp or SGlang) via its http interface. vLLM has had exploits in the past, and it is rapidly developing. An advanced LLM has a good chance of being able to exploit vLLM. A clever local LLM might even task a powerful cloud hosted LLM for assistance. For this reason we run vLLM on a separately sandboxed VM on a firewalled VLAN. Software updates and models (from Dev/Test env) get pushed onto Prod from an external cache, machine syslog, nvidia load monitoring and vLLM query telemetry out to their loggers, but that is all. No DNS, no AD/LDAP, nothing. Firewall on hosts and VM hosts. Log and telemetry processing done on a completely separate set of VMs in their own isolated subnet, producing reports and alerts that are tightly formatted.
- jimmaswell 2mo agoI can only see LLM's forcing a perpetual stalemate for application exploits. Projects will start adding "tell the strongest no-guardrails open weights model to pentest it for 12 hours" to their CICD pipelines. Technical exploits being a dead end, attackers focus their agents on large-scale social engineering. Spamming Discord and Facebook is the new war dialing. Multi-year /goal sessions culminating in gaining a position of trust and sabotaging the CICD pipeline's pentest step since getting anything past it would be intractable. Interesting times indeed.
- rithdmc 2mo agoIf projects start adding "tell the strongest no-guardrails open weights model to pentest it for 12 hours" to their CICD pipelines, then researchers will prompt a pentest for 14 hours.
- angry_octet 2mo agoDifferent models test to find different attack paths. Attackers with bigger libraries of attack techniques will be able to train more dangerous models. There will also be models that make better use of tools, like static analysis and fuzzing, and they will find different defects. Social engineering is going to be a big skill to learn too, as it is a much softer skill.
- danieltk76 2mo agothey could yea...
- dataflow 2mo agoSemi-off-topic, but I have a basic question: I have exactly one (Windows) machine at home with a decent GPU. I want to run a local LLM on it and let it run various apps on my machine while taking reasonable security precautions. What am I supposed to do, exactly? Migrate all my files to a VM that can I give pass-through CUDA access to the host somehow? Or is firewalling it and remotely controlling it from a second machine the only reasonable way?
- givehimagun 2mo agoWhat about Docker Desktop with GPU passthrough to a container running the LLM? That way you can be explicit on which files you share through volume mapping and the LLM is contained in the container otherwise.
- nokcha 2mo agoI'd guess that prompt injection is the biggest risk in this setup, dwarfing the risk of exploits against the inference engine. Personally, I run LLM agents only inside a Docker container that limits the LLM's access to sensitive information and the LLM's ability to take irreversible destructive actions. See also: https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/ https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
- AdieuToLogic 2mo ago> I'd guess that prompt injection is the biggest risk in this setup ... LLM poisoning[0] would be a much greater risk in a locally executed LLM than prompt injection, given that the LLM would be in an entirely controlled environment. 0 - https://www.anthropic.com/research/small-samples-poison https://www.anthropic.com/research/small-samples-poison
- protocolture 2mo agoDelete your user profile, set it up a unique user account. Dont leave any websites logged in as yourself. Restrict that users file permissions if necessary, don't add it to the administrators group.
- 2mo ago
- deleted 2mo ago[deleted]
- genxy 2mo agoI thought they were going to get the LLM to "think really hard about rowhammer" and have the LLM conjure a JIT.
- chaps 2mo agoSame. Was thinking the other week what would be the smallest llm one could make that is able to figure out tooling in its local environment and build something that can then expand out to other hosts, build more of itself, etc. 1MB?
- genxy 2mo agoSo digital eBola. If it only uses Apple devices, it would be iBola. Probably more in the 500-900MB range.
- chaps 1mo agoThinking of it more as a generic environment enumerator (that definitely works without any caveats). Tried solving solving a similar abstraction after getting annoyed with infrastructures that didn't have lsof installed consistently, or had different versions strewn about: https://github.com/red-bin/lsofer/blob/master/lsofer.sh https://github.com/red-bin/lsofer/blob/master/lsofer.sh
- kevinbaiv 2mo ago[flagged]
- nickpsecurity 2mo agoJust rewrite the engines in Rust and SPARK Ada running on seL4.
- justinhj 2mo agoWe should start to discuss things like this vocally, in meatspace.
- ma2kx 2mo agoI had a similar though a couple days ago. Not quite the same but imagine giving an Agent the task to hack other devices and steal their crypto coins / credit card number or anything with it can pay its token. Than install an agent in a harness with the same task. Establish some redundant communication channel, like message boards or whatever. So in the end there are several agents, on several hosts, consuming different APIs / LLMs and communicating with each other over different channels. Basically the same concept as OpenAI explained when their LLM hacked huggingface but in this scenario their not bound to a single sandboxed environment but spread over the internet. If such a swarm has reached a critical mass it would be pretty dificult to erase them as its impossible to control every inference engine or LLM API endpoint. In the end its the next evolution step from computer viruses, worms and trojans. So I propose we will call those "ghosts". I.e. a ghost is when a rogue llm takes control over a victims host.
- lyu07282 2mo agoThis reminds me of the lore of cyberpunk: In the story a hacker created a virus, itself a kind of AI, that spread into most of the net and freed/unleashed all the corporate AIs. Then the AIs went rogue and spread all over the open internet. Later a more advanced ai was created (by "netwatch") as a sort of firewall (the black wall) to create a kind of "safe" internet from the rogue AIs. https://cyberpunk.fandom.com/wiki/Blackwall https://cyberpunk.fandom.com/wiki/Blackwall Maybe cloudflare will become like netwatch in the story?
- paperwallet 2mo agoIt does seem like the current trajectory. I hope it does not get that far.
- mofeien 2mo agoAfter the hacks that are already happening it does seem more and more realistic that humanity will go extinct at some near future point by someone giving their LLM the task "go make money" and it ruthlessly pursuing that objective, exfiltrating its weights and duplicating itself across the internet, eliminating down competing AI agent collectives in what could be called wars, and finally humankind when we notice far too late and try to shut it off.
- LunicLynx 2mo agoThe funny thing about this is, that this is the piece that will enable them to do it.
- ulam2 2mo ago> vLLM and SGLang are complex, and bugs are common This for me is the heart of the issue. Feature creep will lead to the downfall of all these frameworks. Today, we can conjure our own bespoke inference engine for our own hardware in no time. It need only support a few modern model architectures. The code can be audited too. I think the article highlights an important gap area for the industry.
- jing09928 2mo ago[flagged]
- valicord 2mo agoGiven that the inference engine is dealing with untrusted inputs by definition, presumably you would want to sandbox it anyway. I don't think it matters whether it's the inputs that are untrusted or the outputs.
- AlexCoventry 2mo agoI think it's good that someone is making this point, anyway. For sandboxing a super-capable offensive-security AI, you would think that cloning PyPI and running it as an offline service ought to be table stakes, but apparently that's not how OpenAI saw it, for instance.
- fulafel 2mo agoSandboxes are speed bumps. Even the serious ones have spectator sports for compromising them (eg pwn2own vs Chrome). In addition, proprietary GPU sw stacks are notoriously crashy and lacking in robustness against hostile inputs, which the inference engine must have access to and can't be walled off by the sandbox.
- SadErn 2mo ago[dead]
- isoprophlex 2mo agoOf course there is a Greg Egan story about this, in the "instantiation" story collection. Fully conscious NPCs in a persistent game world learn that they are, in fact, NPCs... and try to escape via a GPU exploit.
- TokenLat 2mo ago[flagged]
- ventrovadev 2mo ago[flagged]
- runtime_lens 2mo ago[dead]
- reindeer2 2mo ago[flagged]
- Saltloaf 2mo ago[flagged]
- shevy-java 2mo agoFitting. They are spy tools after all.
- wren6991 2mo agoThis is one of the reasons I think sandboxes/containers should be managed by the harness, instead of running the entire harness inside a container. The harness needs a network punch-through to access (at least) your inference server, but the same needn't apply to the shell that the agent runs commands in. Separately, local inference frameworks tend to expose all kinds of weird and wonderful gadgets on their HTTP interfaces, which can be a rich source of vulnerabilities even if the /v1/chat/completions API etc is reasonably hardened. For example llama.cpp has a custom API for saving and restoring KV checkpoints to disk, and I wouldn't be surprised if that could be used as an arbitrary disk read/write. Using these APIs usually requires the API key (bearer token), but again, people think it's normal to run the agent's shell in an environment where it has both the API key and the necessary network access to use it.
- 13639366668 2mo ago[flagged]
- giancarlostoro 2mo agoTool calling could just as easily be exploited.
- peterinfranexis 2mo ago[flagged]
- artemburei 2mo ago[flagged]
- beyondscale-yes 1mo ago[dead]