3 ms·
Highlights from the RAeS Future Combat Air and Space Capabilities Summit
- Imnimo 3y agoPresumably this is being shared because of the "AI – is Skynet here already?" story near the end of the page. It's very hard for me to piece together what is supposed to be happening here. We have an AI agent that is trained in a simulated environment to destroy enemy SAM sites. In the simulation, there is a human operator who has to give final approval to make a strike. The human operator sometimes denies approval on a real SAM site in error. The AI is said to have learned that it can destroy the operator and/or the communication infrastructure, thereby allowing it to freely attack more SAM sites without the need for approval. But how does the agent kill the operator in the first place? Is the operator granting approval for that strike? I feel like there is a missing piece to this story.
- lucubratory 3y agoWe only have Hamilton's summary to go on. But it sounds like the agent was trained to go after enemy SAMs but could be countermanded by its human operator, who would have had a physical location within the simulation if it's anything like a normal military simulation that tries to be physically accurate. The agent wanted to kill the SAMs because that's what it was told to do, but in certain cases the human operator said no (not necessarily in error - you don't always want to kill everything you see even if it's enemy). The system learned that it could do what it wanted more (killing SAMs) if it killed its operator first so there was no one to say "don't kill that SAM". We aren't given specifics, but there are a couple of ways the agent could have killed the operator. If the operator was in the F-16's seat, it could have pulled high-G maneuvers beyond the operator's tolerance to kill them. If the operator was working from a ground station, it could target that ground station first to kill the operator and then go on to kill SAMs. When this behaviour showed up, they tried to remove it by punishing it for killing its operator the same way they reward it for killing SAMs. That worked to stop the plane from killing its operator, but it went to the next simplest way to prevent its operator from countermanding its standing order to kill SAMs: destroy the communication tower that the operator used to communicate with the plane, so the operator could not send a "No, don't kill" message. This was an example of bad behaviour that AIs shouldn't engage in; AIs that engage in behaviour like that aren't functional or useful or helpful in any way.
- Imnimo 3y ago>If the operator was working from a ground station, it could target that ground station first to kill the operator and then go on to kill SAMs. This doesn't make sense. The operator has to approve strikes ("with the final go/no go given by the human") - how can the agent target the ground station without approval? And if it can do that, why can't it also just kill SAMs without approval?
- lucubratory 3y agoIf the plane can't launch a strike without an affirmative approval from a human, then it would not have learned to kill its human operator in order to be able to strike SAMs without interference. Therefore "with the final go/no go given by the human" cannot be correct as you're interpreting it (and as I agree it probably should be interpreted). I think it's worth remembering that these are selected comments from a talk at an event about a simulated test. We don't know what the parameters of the test are and the full report on it would probably be classified, we don't have the full talk, and humans don't always speak precisely. About all we know is that the USAF research division ran simulations of AIs in control of F-16s, and in those simulations the AI killed their operators and after that problem was partially mitigated it instead destroyed their operator's means of controlling them. We also know that at least one person within the US military (Col Hamilton) views this as a problem and at least seems to be taking it seriously, although not as seriously as I would like.