3 ms·
Couldn't you train the model to keep score (develop a heuristic) for it's own level of certainty for a given answer?
by codebolt 3y ago
Couldn't you train the model to keep score (develop a heuristic) for it's own level of certainty for a given answer?
- blackbear_ 3y agoIt has already been done: http://arxiv.org/abs/2207.05221 http://arxiv.org/abs/2207.05221 However, it seems that RLHF considerably reduces the model's calibration, so perhaps the method above won't be applicable to ChatGPT and similar.