3 ms·
I don't see why this isn't a focus in RLHF. Just heavily discourage confident lies and encourage admission of ignorance. It can't be that much harder than "safe
by Silphendio 3y ago
I don't see why this isn't a focus in RLHF. Just heavily discourage confident lies and encourage admission of ignorance. It can't be that much harder than "safety" training.
But then we would probably end up with things like "As an AI language model, my knowledge of the world is limited, but if I were to guess, ..."