3 ms·
Excellent background on knowledge calibration from Anthropic: https://arxiv.org/abs/2207.05221 https://arxiv.org/abs/2207.05221 "Calibration" in a knowledge c
by xianshou 2y ago
Excellent background on knowledge calibration from Anthropic:
https://arxiv.org/abs/2207.05221 https://arxiv.org/abs/2207.05221
"Calibration" in a knowledge context means having estimated_p(correct) ~ p(correct), and it turns out that LLMs are reasonably good at this. Also a core reason why LLM-as-a-judge works so well: quality evaluation is vastly easier than generation.