4 ms·
I've thought about this before but I think you have normalize the confidence over lots of examples for it to not lead to a degenerate solution. So if the confi
by SomewhatLikely 4y ago
I've thought about this before but I think you have normalize the confidence over lots of examples for it to not lead to a degenerate solution. So if the confidence is divided by the sum of confidence for the whole minibatch the model would have an incentive to spread this out correctly rather than just always hedging.
- nerdponx 4y agoThere are plenty of fancy techniques out there already for building probabilistic neural networks. But I'm not aware of any results that combine them with large language models to develop a confidence score over an entire response. I wonder if people don't even want confidence scores when they say they want machine learning in their product: they want exact answers, and don't want to think about gray areas.