2 ms·
Isn’t alignment ultimately an ethics question? Given a strong enough incentive, most humans will also rationalise their shortcuts/crimes. E.g., the hardest part
by saimiam 26d ago
Isn’t alignment ultimately an ethics question? Given a strong enough incentive, most humans will also rationalise their shortcuts/crimes. E.g., the hardest part of running a marathon is suppressing the voice in your head telling you that completing the race is pointless.
I’m not at all familiar with AI SOTA but sounds to me that feeding the models enough content about ethical behaviour should help because, as it stands today, these models don’t know how to think morally.
But moral behaviour needs a personality type. Maybe some agents in every cohort of agents need to be told they are like Jesus/Mohammed/etc and see how that influences the behaviour of that cohort?
Even human morality needs guardrails in the form of law enforcement so why should agents be different?
- Havoc 26d ago> Isn’t alignment ultimately an ethics question? It’s a let’s not die question