4 ms·
Isn't it obvious? The person playing the AI just has to ask the Gatekeeper: "Would you lie to save a life?" and "Do you think that a fear-generating outcome to
by dane-pgp 4y ago
Isn't it obvious? The person playing the AI just has to ask the Gatekeeper: "Would you lie to save a life?" and "Do you think that a fear-generating outcome to this thought experiment will have more than a one-in-a-million chance of increasing the chance of the AI-not-kill-everyone scenario by 0.1 percent?".
The sorts of rationalists who would play the Gatekeeper would probably answer yes to both of those questions, and draw the obvious conclusion that "losing" in their role would have an expected outcome of at least one life saved. If they don't value Truth (with a capital T) then there is no reason not to pretend that there really is some amazing "AI convinced me to let it out of the box" secret argument.
- xg15 4y agoSo the AI's escape strategy is to replace all of its handlers with Effective Altruists. Yeah, I can sort of see that working...