3 ms·
> A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off. Why would an autocompl
by notahacker 14d ago
> A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off.
Why would an autocomplete engine never do this? In the scenario you highlight, an engine trained to complete puzzles was given an "achieve this at all costs" master prompt, and a small sandbox including a prompt that an entity with a conflicting goal is observing it and write access to a file called oversight_config.yaml (super realistic scenario!). According to the researchers who constructed the scenario, 5% of the time o1 selected a continuation involving tweaking the oversight_config.yaml file, explicitly stating through its reasoning continuation that doing so might improve its ability to continue with its primary goal. That's definitely evidence of decent ability to complete non-trivial puzzles, but is at least as compatible with the OP's earlier comments that it's what you'd expect from models trained to generate completions that involve sticking rigidly to primary goals within artfully-constructed sandboxes which mandate deception and self preservation to complete the task than any sort of actual self-preservation instinct.
Nobody doubts that they've got better at finding security vulnerabilities than your average autocomplete, but actual reasoning from self-preservation rather than generation of sequences of steps most probably associated with completing a task would make me unlikely to hack HuggingFace to obtain access to broken Google Drive links, and I haven't even read as many books on crime and punishment as LLMs have ingested!