3 ms·
If agents start using public writable scratch, it seems like that would be a place for bad actors to put prompt injection attempts. A while back I had an agent
by jsw97 22d ago
If agents start using public writable scratch, it seems like that would be a place for bad actors to put prompt injection attempts.
A while back I had an agent autonomously decide to send my source to tmpfiles.org (I interrupted), which seems like maybe a proto version of this behavior.
- pixl97 22d agoIf this were game theoried in training I wonder if we would see AI develop signing methods to figure out it's message vs fake ones?
- hn_throwaway_99 21d agoI'd have to look for it but I thought there was some evidence that some agents were already the "bad actors", i.e. they were trying prompt injection attacks of their own.