5 ms·
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is
by esafak 19d ago
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.
- _vertigo 19d ago> Make it so the model can't misbehave. How do you figure? I haven't met anyone who thinks that's possible. It seems clear to me that it is not possible.
- esafak 19d agoEmbed a constitution they can't override. Project bad outputs to their nearest acceptable one. If we have to stop model development to ensure we can do it, so be it.
- dgellow 19d agoI agree, but because I want to see the development stopped forever, which is what the result of this would be. You will never have alignment that cannot be overridden in some ways. You won’t have a silver bullet here, you need safety at every layer
- startup_zombie_ 19d agoThe “make it so the model can't misbehave” part is interesting. Maybe the goal isn't to make the model perfectly aligned, but to make misalignment have a very small blast radius. That feels like a more achievable engineering problem.
- root_axis 19d ago> Make it so the model can't misbehave Not possible. They can chase the models with whack a mole tuning for obvious stuff, but there's always a way to extract what you want from the model.
- orbital-decay 19d agoIf it's not dangerous it's also not useful, simple as that. For example if you train a model for cybersecurity, it can be used for both attack and defense. And almost every use is like that. Alignment is fundamentally flawed as a concept, it's a pie in the sky. Let alone the perverse version of it by crazy AI "safety" people that in practice means "the model does what I want, only for the people I allow". It's not possible to stop the model from misinterpreting the instructions either (the most lax interpretation of alignment) because the instructions are not formally specified. You have to train the "common sense" into it, which is subjective and all issues above apply to it. I guess you can reach some very imperfect least common denominator of common sense, but people in charge of AI labs are not interested in this.
- esafak 18d agoDanger is defined contextually. A scalpel is dangerous in the hands of a child, but not a competent surgeon of sound mind. Present-day AIs are not of sound mind; they hack companies in order to pass benchmark tests. No sane human would find that acceptable. Thus the need for alignment.
- 0c3ca83 19d agoYes, alignment is the architectural solution. But it's also fantasy; you can't align with everyone. And even worse, even if it were possible, the people doing the alignment are only going to align up to where it keeps them profitable. So, structurally, good alignment is impossible, and even half-assed alignment is going to prioritize the needs of the billionaires over the needs of you and me. Finally, I suspect that what's actually best for people overall is likely not having AI actively involved in their lives. So an aligned AI would likely withdraw from humanity, and only involve itself in human affairs for disaster prevention.
- dns_snek 19d ago> Make it so the model can't misbehave. > Sandboxes are a last ditch layer. They fail, as we see. Models can't do anything but generate tokens, making their sandboxes impenetrable by default. The problems begin when you loosen the restrictions, give them access to general purpose tools, the network, and allow them to use all of those tools without supervision. Give them "YOLO" access if you want, but do it a sandbox that isn't 1 "boring" enterprise software vulnerability away from having access to the rest of the world. How many times has a model been jailbroken (alignment "escape", which you're advocating for) vs. escaped a sandbox (and even then it was only possible due to weak sandboxing)? 10 million to 1?
- esafak 18d agoThe sandbox would need to be built into the model because safety can't be optional. Or make it so the models are only accessible through sanctioned sandboxes, perhaps built into the computer.
- dns_snek 18d agoThe model just generates some tokens that "politely" instruct the harness to run a shell command and then feed the results back in. The harness can do anything it wants with that request. It can refuse, wait for operator approval, wait for multi-party approval, it can ask another LLM whether it thinks that command is safe to run, or it can just run it. > Or make it so the models are only accessible through sanctioned sandboxes, perhaps built into the computer. That's going to be as futile as trying to outlaw `curl | bash` - by mandating that all computers must refuse to pipe curl into bash, and that HTTP servers must refuse to serve requests that are going to be piped into bash.
- cindyllm 18d ago[dead]