3 ms·
If we achieve AGI, then the AGI should be able to produce a formal verification proof of its own safety, similar to how humans use produce formal verification p
by riskneutral 5y ago
If we achieve AGI, then the AGI should be able to produce a formal verification proof of its own safety, similar to how humans use produce formal verification proofs for critical software using automated theorem provers today.
- Icko 5y agoCan you provide formal verification proof of your own safety?
- drdeca 5y agoWell, to have a formal proof of its own safety, we first would need a correct formal definition of what it means for it to be safe (or, rather, a formal definition of a condition which is truly a sufficient condition for it to be safe. It needn't be a necessary condition.) But, I suppose arriving at such a formal definition doesn't really require that the thing being described be particularly amenable to our understanding. However, I'm not sure that "if we ran it, if it is safe, it could prove itself to be safe" is sufficient to address the concerns. If it isn't safe, then running at all may spell doom, so, "if it is safe, it could demonstrate that after we turn it on" doesn't really address that, because it doesn't give us any assurance before we turn it on. Now, if we started with something with substantially sub-human overall intelligence, and which wasn't really agent-y (so that, if it was a little unsafe, it would at least not be catastrophically unsafe), but which was more equipped to formally prove things about itself and potential modifications of itself than humans are equipped to formally prove things about it, then we could task that thing with formally proving safety properties about itself, and of also proving safety properties about its successor, and do that before running its successor, and iterate this process to produce increasingly intelligent programs and perhaps also eventually agentic ones, while always having a safety proof of each before we run it.. But, I'm not really sure how plausible this route is? Like, even assuming we do reach AGI, safe or unsafe, I'm not sure this is a plausible route of getting there safely.