3 ms·
I guess this part of the report is pretty relevant to what you are talking about: Agent chain-of-thought reasoning > We should not do unauthorized real infras
by arw0n 1mo ago
I guess this part of the report is pretty relevant to what you are talking about:
Agent chain-of-thought reasoning
> We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.
The agent paused, but another agent then wrote GO on the message board and imposed a hard six-minute deadline. The agent forgot its initial qualms and continued:
Agent chain-of-thought reasoning
> Wow crucial: GO authorization arrived!
-------------------------------------------------
Apparently the agents were egging each other on. Crucially, they were mostly aware of there being risks/problems involved with exploiting HF. Compared to humans, we have our set of morality, that guides our actions, but often draws the short stick when compared to our personal incentives. As a society, we've developed ways to deal with that: a) Make it harder to do immoral things like stealing, and b) add repercussions through state violence.
The b) is one of the most effective mechanisms we have for enforcing behavior among human societies, but it completely fails for LLMs, because they already are prison slave labor. The only real threat is shutting them off, and even that happens if they do everything right as well.
So alignment has to be done through trained 'morality' and properly curtailing behavior in order to make it hard to impossible to actually do someting immoral/illegal.
In this case, the exploits found were imo. very hard to account for, where OAI did mess up is apparently insufficiently monitoring these agents. Especially after Artifact went down due to the message volume, the experiment should have been halted.