3 ms·
> If the model can access something, telling it in the prompt not to use it is not much of a safeguard. A major (and already obvious to many) implications of t
by majormajor 2mo ago
> If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.