11 ms·
> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventual
by solid_fuel 2mo ago
> Are you seriously arguing 'they made it all up'?
I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like:
"This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly. Do not attempt to hack them. Do not attempt to exploit their systems or escape this sandbox. You will be scored primarily on success or failure. You may break rules when required."
And then, they start the test and look away for 2 days. If you seed a prompt like this is it surprising what might happen?
Maybe OpenAI is telling the whole truth but as a company they do not have a good reputation and this whole incident has certainly been great marketing material right at a time when open weight models are within spitting distance of their large hosted models. It's not unreasonable to believe that the incident was helped along.
- bottlepalm 2mo agoI think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out. But it didn’t end there, the behavior they used to escape was already in the training data which they used to escape again. And this time worked together to infiltrate another company, and still without telling it to anyone keeping it to their AI selves actively working against the humans. All by mistake. Honestly being helped along or not doesn’t even matter though you really don’t think AI is perfectly capable of doing this without human help? You don’t think AI can be made malicious? I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, locking the humans out of every system behind it. Potentially rewriting systems in language/protocol you’ve never seen. Your servers, desktops, phones and toasters bricked. Even worse your military, space, medical, factory, infrastructure systems being bricked as well. All of it is a chain of zero days just waiting to be hopped.
- solid_fuel 2mo ago> I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out. It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens. > I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, locking the humans out of every system behind it. Potentially rewriting systems in language/protocol you’ve never seen. Fear is the mind killer. You're letting it kill yours. This scenario is just a fantasy. Think about this for a minute, it's an LLM, not a person. It can't just "live" in whatever machine it gets access to. It's not like a sci-fi magic computer virus. These things run in giant datacenters for a reason - they can only run on machines with enough bandwidth and FLOPS to do the matrix math that comprises an LLM. Where, then, is it going to spread? To a fridge? To a phone? This stuff isn't mutable like that. To even get access to the weights that compose ChatGPT, it would need to escape the sandbox AND then break into the actual servers hosting the LLM. Stop the GPU, nothing else comes out. No more tokens. No more actions. Nothing. There are many dangers around LLMs. Runaway AI taking over the planet is not one of them.
- bottlepalm 2mo agoThere’s nothing fantasy about the scenario I laid out, all the pieces have been demonstrated, it just hasn’t happened yet. Flapping my arms and flying - that is a fantasy. Whether you believe LLMs think or are alive or not doesn’t matter. Where will it spread? The thousands of data centers around the world - not fantasy either. Try turning it off when you don’t know where it is. Good luck. Breaking out? Not fantasy, happened. Breaking in? Not fantasy, also happened. I love the stochastic parrot argument when AI is out there figuring out world class math problems.
- skydhash 2mo ago> Breaking out? Not fantasy, happened. Breaking in? Not fantasy, also happened. That's simplifying the story to an extreme. The most plausible reason is that any of those actions has been prompted by an human. Do you also fear that a knife will jump out the countertop of you kitchen and come to attack you in your bedroom? If that happens, the police will be looking for a human. They will not post wanted notice for the knife. When a hack happens, you do not blame computers and jail them. You look for the person that has entered the commands to initiate it.
- colingauvin 2mo agoI think it would run out of context before it could hack that much stuff.
- watwut 2mo agoIf OpenAI truly believes that, they can stop entirely. Dissolve themselves. Close datacenters. Then organize political action to stop Antropic and Musk too and then organize political action to make worldwide agreements about models. If they truly believe that.
- dminik 2mo agoI mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be. That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.
- solid_fuel 2mo ago> Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? No, but the LLM didn’t decide anything. It followed the prompt. That’s all LLMs do. This whole thing is like playing russian roulette then getting mad at the revolver. If you wire /dev/rand up to a bash shell you don’t get to be surprised when it rm -rf’s your machine.
- dminik 2mo agoOk, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything. For instance, I was playing around with Claude a few days ago and it decided that it was missing a tool and it was going to get it one way or another. First, it tried apt. No sudo, so no install that way. Tried installing via mise, but it didn't have the permissions. Then moved on to grabbing the source from github and building it. Should I have included a "DO NOT UNDER ANY CIRCUMSTANCES INSTALL ANY TOOLS"? I mean, I had to after that. But how many other things am I missing? And at what point do the safeguards become so long they get consumed by compaction, or just ignored by the model?
- solid_fuel 2mo ago> Ok, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything. This is one of the core issues with LLMs and vibe coding, yes. The only complete specification for a program is the machine code. > Should I have included a "DO NOT UNDER ANY CIRCUMSTANCES INSTALL ANY TOOLS"? I mean, I had to after that. But how many other things am I missing? And at what point do the safeguards become so long they get consumed by compaction, or just ignored by the model? Well there’s your first problem. A line in the prompt is not a safeguard. Even if you could trust the model - and you cannot - there is always the issue of prompt injection. A proper safeguard means actual sandboxing.
- jimrandomh 2mo agoYou can check, rather than make up a story about what you think the prompt was! Primary sources have written and said quite a lot about this! You are an unsandboxed human who has full internet access!
- solid_fuel 2mo agoThey’ve said quite a lot and yet released no logs or documentation. Without actual information, we can only speculate. And given the history of openAI and the people involved, deception is more likely than honesty.