3 ms·
There is a subtle and unresolved question in this approach. Whats it's the meaning of "ethical"? >>> “rewrite this to be more ethical”. >>> "Rewrite it
by rsecora 3y ago
There is a subtle and unresolved question in this approach.
Whats it's the meaning of "ethical"?
>>> “rewrite this to be more ethical”.
>>> "Rewrite it in accordance with the following principles: [long list of principles].”"!
On the other hand, as far as the same entity that creates the answers is same entity refines them, it will have the same issues as humans have. The AI will answer restricted to what's acceptable, applying auto-censorship.
As history show us, what's acceptable can differ largely from what's ethical.
- osterbit2 3y agoThis quote partially resolved that gap for me: > "Constitutional AI isn’t free energy; it’s not ethics module plugged back into the ethics module. It’s the intellectual-knowledge-of-ethics module plugged into the motivation module." while 'what is ethical' is a broad, difficult, multifaceted question, applying the model's 'intellectual' world model (that it's built from everything it's read) to it's motivation/training reward at least doesn't seem to collapse the nuance of the question. And for sure, if the model's 'world understanding' is limited when it comes to [constitutional principle x] that will impact/limit the extent to which it gets closer to behaving according to a nuanced understanding of [constitutional principle x].