3 ms·
Great work! I found the results in Fig 16 pretty interesting [0]... From response 1, it seems that the model has very little confidence in its decision but get
by muds 5y ago
Great work! I found the results in Fig 16 pretty interesting [0]...
From response 1, it seems that the model has very little confidence in its decision but get the correct answer while, on the contrary, in response 3, the model seems very confident in its incorrect answer. Is this usually a trend that you see with large models? How hard is it, generally, to make such models "aware" of their own shortcomings?
[0]: https://arxiv.org/pdf/2108.07732.pdf#figure.caption.21 https://arxiv.org/pdf/2108.07732.pdf#figure.caption.21
- tasdfqwer0897 5y agoI think it might be a mistake to think that the model is not confident because its response is something a human might say if they were not confident. The model is 'just' completing the prefix text with something that has high likelihood from its perspective, so it may just be used to, for instance, seeing people hedge in similar conversations it has read in its training data. More generally, whether these models are well-calibrated (that is, they know what they don't know) is an important area of research. I don't have references offhand, but I think it's true broadly speaking that these larger pre-trained models do tend to be better calibrated.