4 ms·
Thinking through the concept, I imagine that if the LLM was being used as an agent to execute commands and the attacker knew the tools it would interface with,
by batch12 3y ago
Thinking through the concept, I imagine that if the LLM was being used as an agent to execute commands and the attacker knew the tools it would interface with, this maybe would be possible. More interesting to me would be crafting an exploit for the underlying python/cpp being used to run inference and training the model on this. Then maybe drop a trigger which would generate the exploit and allow the execution of additional code. Now, maybe this isn't feasible through training. Maybe some clever payloads could be crafted and pulled in by the model during RAG to do this which seems like a more plausible method of attack to me.