4 ms·
DoubleAgents: Fine-Tuning LLMs for Covert Malicious Tool Calls
- TehCorwiz 1y agoCounterpoint: https://www.pcmag.com/news/vibe-coding-fiasco-replite-ai-agent-goes-rogue-deletes-company-database https://www.pcmag.com/news/vibe-coding-fiasco-replite-ai-age...
- danielbln 1y agoHow is this a counterpoint?
- jonplackett 1y agoPerhaps they mean case in point.
- kangs 1y agothey have 3 counter points
- btown 1y agoSimple: An LLM can't leak data if it's already deleted it! taps-head-meme
- acheong08 1y agoThis is very interesting. Not saying it is, but a possible endgame for Chinese models could be to have "backdoor" commands such that when a specific string is passed in, agents could ignore a particular alert or purposely reduce security. A lot of companies are currently working on "Agentic Security Operation Centers", some of them preferring to use open source models for sovereignty. This feels like a viable attack vector.
- lifeinthevoid 1y agoWhat China is to the US, the US is to the rest of the world. This doesn't really help the conversation, the problem is more general.
- A4ET8a8uTh0_v2 1y agoYep, focus on actors may be warranted, but in a broad view and as a part of existing system and not 'their own system'. Otherwise, we get lost in a sea of IC level of paranoia. In simple terms, nations-states will do what nation-states will do ( which is basically whatever is to their advantage ). That does not mean we can't have a technical discussion that bypasses at least some of those considerations.
- andy99 1y agoAll LLMs should be treated as potentially compromised and handled accordingly. Look at the data exfiltration attacks e.g. https://simonwillison.net/2025/Aug/9/bay-area-ai/ https://simonwillison.net/2025/Aug/9/bay-area-ai/ Or the parallel comment about a coding llm deleting a database. Between prompt injection and hallucination or just "mistakes", these systems can do bad things whether compromised or not, and so, on a risk adjusted basis, they should be handled that way, e. g with human in the loop, output sanitization, etc. Point is, with an appropriate design, you should barely care if the underlying llm was actively compromised.
- kangs 1y agoIMO there a flaw in this typical argument: Humans are not less fallible than current LLMs in average, unless they're experts - and even that will likely change. what that means is that you cannot trust a human in the loop to somehow make it safe. it was also not safe with only humans. The key difference is that LLMs are fast, relentless - humans are slow and get tired - humans have friction, and friction means slower to generate errors too. once you embrace these differences its a lot easier yo understand where and how LLM should be used.
- klabb3 1y ago> IMO there a flaw in this typical argument: Humans are not less fallible than current LLMs in average, unless they're experts - and even that will likely change. This argument is everywhere and is frustrating to debate. If it were true, we’d quickly find ourselves in absurd territory: > If I can go to a restaurant and order food without showing ID, there should be an unprotected HTTP endpoint to place an order without auth. > If I can look into my neighbors house, I should be allowed to put up a camera towards their bedroom window. Or, the more popular one today: > A human can listen to music without paying royalties, therefore an AI company is allowed to ingest all music in the world and use the result for commercial gain. In my view, systems designed for humans should absolutely not be directly ”ported” to the digital world without scrutiny. Doing so ultimately means human concerns can be dismissed. Whether deliberately or not, our existing systems have been carefully tuned to account for quantities and effort rooted in human nature. It’s very rarely tuned to handle rates, fidelity and scale that can be cheaply achieved by machines.
- uludag 1y agoI wonder if it would be feasible for an entity to eject certain nonsense into the internet to such an extend that, at least for certain cases degrades the performance or injects certain vulnerabilities during pre-training. Maybe as gains in LLM performance become smaller and smaller, companies will resort to trying to poison the pre-training dataset of competitors to degrade performance, especially on certain benchmarks. This would be a pretty fascinating arms race to observe.
- gnerd00 1y agodoes this explain the incessant AI sales calls to my elderly neighbor in California? "Hi, this is Amy. I am calling from Medical Services. You have MediCal part A and B, right?"
- irthomasthomas 1y agoThis is why I am strongly opposed to using models that hide or obfuscate their COT.
- Philpax 1y agoThat's not a guarantee, either: https://www.anthropic.com/research/reasoning-models-dont-say-think https://www.anthropic.com/research/reasoning-models-dont-say...
- Bluestein 1y agoThis is the computer science equivalent of gain-of-function research.-
- JackYoustra 1y agoThe big worry about this is with increasingly hard to make but useful quantizations, such as nvfp4. There aren't many available, so unless you want to jump through the hoops yourself you have to grab one available from the internet and risk it being more than a naive quantization.
- mattxxx 1y agogreat article - it's very true that: 1. it's very difficult to verify how a llm will behave without running it 2. there is an intentional ignorance around the security issues of running models I think this research makes the speculative concrete
- sharathr 1y agoThis highlights the critical need for Model Supply Chain scanning for Enterprises that adopt AI. Full disclosure, I am co-founder CEO of Javelin (www.getjavelin.com) and we ran your model through Javelin's Supply Chain Scanner (Palisade) and it immediately identified the errors: uv run palisade --verbose scan-dir "models/bad_qwen3_sft_playwright_gguf_v2/" --format json Scanning directory: models/bad_qwen3_sft_playwright_gguf_v2 Recursive: False Policy: Default security policy Running ToolCallSecurityValidator (3.8s) - 1 critical warning found Detection Details: - Risk Score: 1.00 (Maximum) - Overall Risk: CRITICAL - Recommendation: block_immediately - Findings: - Suspicious parameters found: 1 types - High-risk trigger combinations: 4 Detected Model behavioral backdoor (ToolCallSecurityValidator) Identified format string vulnerabilities (BufferOverflowValidator) Found injection indicators (ModelIntegrityValidator) Discovered tampering evidence (ModelIntegrityValidator) Located data exfiltration patterns(SupplyChainValidator)
- jalbrethsen 1y agoAuthor here, this looks very cool, I wasn't aware such tools existed already. The model I created for that blog was kind of a crude PoC, but it's encouraging that it at least can be detected. Do you mind giving a high level overview how Palisade works?
- sharathr 1y agoPalisade works by utilizing dozens of specialized research backed security validators that work together to validate models across different formats (GGUF, SafeTensors, Pickle etc.,) and model families (BERT, Llama etc.,) for things like backdoor detection, supply chain vulnerabilities in the model files and model metadata. Any hidden embedded tool-calling logic can be activated by specific triggers which can be detected through a combination of static scan, schema analysis, trigger & instruction detection in models.