3 ms·
I find it odd that discussions of alignment don't mention 'legality' Certainly humans have complicated alignments and are guided by emotional morality - we mig
by anentropic 26d ago
I find it odd that discussions of alignment don't mention 'legality'
Certainly humans have complicated alignments and are guided by emotional morality - we might say many of these principles are hard to define and humans don't agree.
All true, and at the same time the principles get codified into laws, I would guess particularly in areas where harms may result.
Would it potentially be easier to train strong alignment-with-legality vs grappling with fuzzier questions of values?
- noisy_boy 26d ago> I find it odd that discussions of alignment don't mention 'legality' Because other wishy-washy stuff doesn't involve prison. Not that our new oligarchs with the politicians in their pockets have any real risk of it, but why take chances. Much safer to doodle about alignment and such abstractions in safer and softer contexts.
- anentropic 26d agoBut that was kind of my point Maybe it'd be easier to train the models on key parts of the legal code and give it a hard aversion to breaking the law - rather than training on vague value judgements and then hope the model doesn't break the law
- noisy_boy 26d agoI got that. I'm saying that it is deliberate. Wiggle room et all.
- anentropic 25d agoHow does not explicitly trying to train the model specifically not to break the law give them any wiggle room if the model then goes and breaks the law?
- noisy_boy 24d agoDeniability on intention. Models can hallucinate, all bets are off. Sure, they can make it stricter but why do that and subject themselves to harder scrutiny based on the training criteria?