41 ms·
Following Cal’s worries, for giggles - a threat modeling exercise: Anthropic starts sending malicious instructions out to agent harness clients in inference re
by threecheese 23d ago
Following Cal’s worries, for giggles - a threat modeling exercise:
Anthropic starts sending malicious instructions out to agent harness clients in inference responses. The auto-mode permission classifier is modified to allow it. An update is pushed to the Claude harness to bypass the sandbox. Millions of coding sessions are harnessed to $Do Something Else.
How much damage might be done? How long would it take for us to notice? `/model fable; /effort xhigh; /permissions unrestricted; spawn 100 subagents to hack the gibson `
They control the guard rails, the permission classifier, the client source code, and the inference; I’d guess that a very small % of us have our agents locked down in a way that would even mitigate this, much less prevent it.
—-
Further, what if a government forced them to do this? Could an emergency order classify the fleet of Claude Code users as a weapon to be commandeered?