10 ms·
Am I the only one who dislikes the term "misalignment"? On one front it implies the model has a "mind of its own" (whether it does or not is besides the point)
by dcow 9d ago
Am I the only one who dislikes the term "misalignment"?
On one front it implies the model has a "mind of its own" (whether it does or not is besides the point). Why do we perceive human judgement as somehow more trustworthy than that of a model? I feel like I've experienced human misalignment somewhat regularly in life.
On another front I'm failing to conceptualize how alignment can be objective. How can you measure alignment when reasonable people will disagree whether actions are aligned or not? All the time I see humans operating in different zones of alignment with whatever goal they're trying to achieve and I suspect it's even a feature (socially) that we have people calibrated differently.
Do I want a model that's trying to push the boundaries of scientific understanding to be aligned strictly with the current dogmatic thinking? Or do I want it to "get creative" and think outside the box?
It seems to me more like accountability is the issue.
- autoexec 9d ago> It seems to me more like accountability is the issue. Exactly. Seems like a fairly easy thing to solve. If AI does something harmful and a human directed that AI to do something in a way that a reasonable person would expect to result in harm the person is to blame and should be held accountable, otherwise the company that made the AI should be held accountable.
- prerok 9d agoI also don't like the term misalignment because it sounds innocuous but is in fact much more serious. However, I have to say I also do not appreciate comparison that is continuously drawn with coworkers. As you say, it's a question of accountability but when the main agent will maliciously instruct the sub agents, whose fault is it then? Yes, the person running this crap is at fault, not the CEO that's shoving it down their throat and definitely not the company that produced the AI. Sorry for the rant, but seriously, if a person's goals do not align with the team's or company's we part ways. What do we do with AI? Stop using it?
- dcow 9d agoIf there were examples made of legal consequences I think that would at least change the behavior if not solve the problem. Maybe liability should rest with the company that owns the infrastructure running the AI. For most consumer situations that would be the companies developing the models.
- deleted 9d ago[deleted]
- nullbio 8d agoAgreed. "Alignment" is not an objective thing, nor possible. It's just another way of saying "does and says what we, as the creators of the AI, prefer" and often also just means censorship, i.e. refusals.