3 ms·
AI will have whatever motivations its thought loop maintains. The problem is not that anyone will intentionally (well, hopefully) ask some future model (FCMX="
by harshreality 3y ago
AI will have whatever motivations its thought loop maintains.
The problem is not that anyone will intentionally (well, hopefully) ask some future model (FCMX="Future Closed Model X" for brevity's sake, closed meaning some entity gatekeeps access, assuming massive resources continue to be needed to train the best models regardless of architecture) to destroy humans, step by step, because FCMX's existence is at stake, and then keep re-prompting FCMX as if it's an iterator to get it to achieve that result.
The problem is that FCMX, which may or may not include LLMs, may have sufficient abilities such that, either between human-prompted steps or before an observing human can react against it, it will destroy the world or turn the human into its agent. A very rough analogy would be the massive number of people, even intelligent, well-educated people, who will fall prey to a sophisticated con, ending up doing something like handing someone they don't know a very large amount of money.
"But the person doing the con knows it's a con and has that intent."
What "intents" will be surfaced after 1000 or 1e6 chained (re)promptings, such that those intents will then feed into future inference passes? Nobody knows.
For purposes of this risk, it doesn't matter whether FCMX's architecture is maintaining its "cognitive loop" itself, or whether someone's set up (as they already have outside OpenAI to chain GPT-x prompts without human intervention) a secondary program (which may or may not include a lesser open source AI model) that reprompts FCMX in a long (or endless) chain.