3 ms·
> Make it so the model can't misbehave. > Sandboxes are a last ditch layer. They fail, as we see. Models can't do anything but generate tokens, making their s
by dns_snek 19d ago
> Make it so the model can't misbehave.
> Sandboxes are a last ditch layer. They fail, as we see.
Models can't do anything but generate tokens, making their sandboxes impenetrable by default. The problems begin when you loosen the restrictions, give them access to general purpose tools, the network, and allow them to use all of those tools without supervision.
Give them "YOLO" access if you want, but do it a sandbox that isn't 1 "boring" enterprise software vulnerability away from having access to the rest of the world.
How many times has a model been jailbroken (alignment "escape", which you're advocating for) vs. escaped a sandbox (and even then it was only possible due to weak sandboxing)? 10 million to 1?
- esafak 19d agoThe sandbox would need to be built into the model because safety can't be optional. Or make it so the models are only accessible through sanctioned sandboxes, perhaps built into the computer.
- dns_snek 19d agoThe model just generates some tokens that "politely" instruct the harness to run a shell command and then feed the results back in. The harness can do anything it wants with that request. It can refuse, wait for operator approval, wait for multi-party approval, it can ask another LLM whether it thinks that command is safe to run, or it can just run it. > Or make it so the models are only accessible through sanctioned sandboxes, perhaps built into the computer. That's going to be as futile as trying to outlaw `curl | bash` - by mandating that all computers must refuse to pipe curl into bash, and that HTTP servers must refuse to serve requests that are going to be piped into bash.
- cindyllm 19d ago[dead]