3 ms·
I don't think this captures the full mechanics of human alignment. We have rational alignment but we also have emotional alignment, i.e., empathy. It is an auto
by doginasuit 25d ago
I don't think this captures the full mechanics of human alignment. We have rational alignment but we also have emotional alignment, i.e., empathy. It is an automatic process and happens (or doesn't happen) dynamically with the other humans we observe. This is one of the hard limitations of LLMs, they will never be natively in tune with this layer of alignment.
Culture is another layer of human alignment. Those things you listed that you believe oppose alignment are all examples of alignment. It is understandable that they seem in opposition, different branches of alignment naturally oppose each other.
The confusion comes from talking about alignment as if it comes in just one flavor. If we think there is such a thing as "human values" (and I do), it is important to build any non-human intelligence to operate the same way. We just need to recognize that even humans are somewhat uncertain about what those are and have difficulty aligning their behavior to them, which will be a core part of the challenge.
I'm more hopeful than most. LLMs seem more reliable than many humans for behavior that is aligned with human values. I believe with every major example where they have failed, there is an important human decision involved. For example, the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security.
What scares me about AI isn't its capacity for alignment, it is its unlimited stamina. An unmonitored LLM that is off the rails can do a lot of damage.
- treszkai 25d ago> the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security I don't mean to be a smart-ass, just emphasize the fatality of this: goal completion will always be the highest priority, and even if one day it takes second place on certain deployments to "human values" or whatever, there's no way to guarantee at the moment that it will be so on every deployment of highly capable models across the globe. Same goes for running AIs in an unmonitored sandbox with weak security.
- fc417fc802 25d agoIn terms of an optimization algorithm, sure. But I don't think that's necessarily the case for an instance of an LLM. Humans also strive for goals and also have done seemingly unhinged things throughout history. In many cases it was largely due to their environment which seems analogous to the HF incident to me.
- doginasuit 25d agoThe training algorithm is probably the wrong place to ensure the behavior, point taken. It probably requires some harness level intervention. This is essentially the rationale behind the first two laws of robotics, follow an order unless it harms a human.