3 ms·
> That makes sense. No, it doesn't if the "which one do you prefer" was just a prompt continuation task. LLMs can't see their own evaluation of a given text. I
by bmacho 1mo ago
> That makes sense.
No, it doesn't if the "which one do you prefer" was just a prompt continuation task. LLMs can't see their own evaluation of a given text. If you ask them to continue
Which one you prefer
> option 1: human text
> option 2: ai text
And they continue with
option 2: ai text is the better one because [reasons]
then it is not because they evaluated these 2 texts on themselves, observed the evaluation numbers and reported which one is better.
Also, if you instruct humans to come up with the best text they can, and you show them an even better text, they will prefer the better one written by someone else.