2 ms·
This could just be a real-world example of the difficulty of the alignment problem. Having no personal insight into the minds of Google engineers who built thi
by hackerlight 3y ago
This could just be a real-world example of the difficulty of the alignment problem.
Having no personal insight into the minds of Google engineers who built this thing, my assumption is they wanted Gemini to give diverse results when someone asked for a "person" or an "accountant" (reasonable), but didn't think of cases like "English kings" (unreasonable) where the added context changes the distribution. So they added some one-line hack to the system prompt (we know this is how OpenAI achieves this) which unbiased the former distribution but added lots of bias the latter. Quite an easy oversight to make.
This is what AI risk people have been saying. You can't get AI to behave how you want exactly because it's extremely difficult to encode your values in a way where there's no unintended consequences.