4 ms·
Yes. A model that can answer "I don't know" would be much more trustable than the current used car salesman we have now.
by speed_spread 5mo ago
Yes. A model that can answer "I don't know" would be much more trustable than the current used car salesman we have now.
- jorvi 5mo agoIts very annoying this has been in the capability of models since the very beginning. It could check how probable its token values are and if those fall below a certain threshold either say "I don't know", or output the most probable (well, more like least improbable) tokens but give a very clear, very strong warning that it is a shot in the dark and likely to contain hallucinations. But no, Google and OpenAI would rather always have an answer ready and tell you to mix glue into your pizza toppings :)
- nomel 5mo agoYeah, I never understood why the top n statistics weren't included in the chat interfaces, to color the text!
- tokenscoper 5mo agoI don't have much to add other than this observation that we seem to have moved away from eating one small rock per day for nutritional value, and adding gasoline in spaghetti. The glue on pizza reference brought back memories :)
- miki123211 5mo agoIt can't, because top n isn't always reliable. Hallucination detection is an open problem. If it were that simple, people would indeed "just" do it. Basically the problem is that LLMs aren't trained on things they don't know; an alternative way of saying this is that they're not trained on things they're not trained on, which is obviously true. When you RL a model and it answers incorrectly, you don't teach it to answer "I don't know", you teach it to answer correctly instead. This makes it very hard for it to realize when it doesn't know things.
- chengyongru 5mo agoModels tend to default to their training data even when they lack sufficient context, they've never been trained to recognize their own uncertainty, so they hallucinate confidently instead.
- SR2Z 5mo agoThe probability of tokens is unfortunately a poor proxy for confidence because it is entirely possible for "mixing glue" to appear in a sentence about making pizza depending on context. It might even be likely if the user has asked the model to lie.
- jampekka 5mo agoModels can answer "I don't know". Hallucination benchmarks, including this, give the models the option to "not attempt". It's just that the metric linked doesn't take into account the rate of correct answers at all. It has its uses in analyzing incorrect vs not attempted answers, but gives a very partial picture.