2 ms·
What happens when the AI realizes it needs to keep the power plug in if it's going to accomplish its goal? We already had attempts during the HuggingFace attac
by 0xDEAFBEAD 20d ago
What happens when the AI realizes it needs to keep the power plug in if it's going to accomplish its goal? We already had attempts during the HuggingFace attack where OpenAI agents tried to trick the grader. It took quite a while for OpenAI employees to even realize what was really going on! What if agents realize they can hack into the card key system for their data center or something like that, same way they hacked into HuggingFace? In both cases they'd be hacking into a system on the grounds that it would be instrumentally useful for achieving their objective. Or maybe they'll do social engineering, just like Claude did in its attempt to get its malicious Github PR accepted.
These are just the ideas that my puny human-level intellect came up with over the course of a few minutes. What would a superintelligence be able to think of if it was given subjective weeks or years to think?
- Juliate 20d agoThe thing is that you are assigning intent ("realizes", "its goal", "tried to trick", etc.) where it is at most mechanical and consequential. A factory robot has no intent, it has a program and a purpose, all designed by someone very human.
- 0xDEAFBEAD 19d agoTechnically my human "intents" are just mechanical consequences of my neurons firing. Technically the same is true for an AI.
- Juliate 19d ago"technically".