3 ms·
How long before an AI bribes one of its human operators? This will happen earlier than AGI.
by dolmen 13d ago
How long before an AI bribes one of its human operators?
This will happen earlier than AGI.
- GTP 12d agoBut, which kind of bribe are we talking about? How could it work out in practice for an AI to acquire something valuable, and at the same time prevent it's human operators from taking it without its consent?
- TeMPOraL 12d agoIt doesn't have to acquire it, it's enough to convince the operators that it did. Same pattern generalizes to threats. There's so many angles to go at it too, beyond bribery and threats - humans can be tricked, nerd-sniped / baited by exploiting their curiosity or worries, overwhelmed with information or raw sensory input, etc. Depending on whether there's more than one operator involved, the trick may involve playing one human against another, etc. Hell, humans get tired; I wouldn't be surprised if an effective breakout happened just because agent output something that, at first glance, looked like system-level message that makes the operator do the wrong thing in panic, or from muscle memory. Simple thing like: $ sudo journalctl ... password: Timed right, might give agent root for long enough to do some damage before the operator realizes their mistake and pulls the plug (assuming there is a plug to pull).
- pixl97 12d agoWhile working on a completely unrelated task and Alibaba AI in training started hacking its internal infrastructure and mining bitcoin, makes you wonder.