3 ms·
Maybe it requires understanding, maybe there are other ways to get to 'I don't know'. There was a paper posted on HN a few weeks ago that tested LLMs on medical
by t_mann 3y ago
Maybe it requires understanding, maybe there are other ways to get to 'I don't know'. There was a paper posted on HN a few weeks ago that tested LLMs on medical exams, and one interesting thing that they found was that on questions where the LLM was wrong (confidently, as usual), the answer was highly volatile with respect to some prompt or temperature or other parameters. So this might show a way for getting to 'I don't know' by just comparing the answers over a few slightly fuzzied prompt variations, and just ask it to create an 'I don't know' answer (maybe with a summary of the various responses) if they differ too much. This is more of a crutch, I'll admit, arguably the LLM (or neither of the experts, or however you set it up concretely) hasn't learnt to say 'I don't know', but it might be a good enough solution in practice. And maybe you can then use that setup to generate training examples to teach 'I don't know' to an actual model (so basically fine-tuning a model to learn its own knowledge boundary).
- andsoitis 3y ago> Maybe it requires understanding, maybe there are other ways to get to 'I don't know'. > This is more of a crutch, I'll admit, arguably the LLM (or neither of the experts, or however you set it up concretely) hasn't learnt to say 'I don't know', but it might be a good enough solution in practice. And maybe you can then use that setup to generate training examples to teach 'I don't know' to an actual model (so basically fine-tuning a model to learn its own knowledge boundary). When humans say "I know" it is often not narrowly based on "book knowledge or what I've heard from other people". Humans are able to say "I know" or "I don't know" using a range of tools like self-awareness, knowledge of a subject, experience, common sense, speculation, wisdom, etc.
- t_mann 3y agoOk, but LLMs are just tools, and I'm just asking how a tool can be made more useful. It doesn't really matter why an LLM tells you to go look elsewhere, it's simply more useful if it does than if it hallucinates. And usefulness isn't binary, getting the error rate down is also an improvement.
- andsoitis 3y ago> Ok, but LLMs are just tools, and I'm just asking how a tool can be made more useful. I think I know what you're after (notice my self-awareness to qualify what I say I know): that the tool's output can be relied upon without applying layers of human judgement (critical thinking, logical reasoning, common sense, skepticism, expert knowledge, wisdom, etc.) There are a number of boulders in that path of clarity. One of the most obvious boulders is that for an LLM the inputs and patterns that act on the input are themselves not guaranteed to be infallible. Not only in practive, but also in principle: the human mind (notice this expression doesn't refer to a thing you can point to) has come to understand that understanding is provisional, incomplete, a process. So while I agree with you that we can and should improve the accuracy of the output of these tools given assumptions we make about the tools humans use to prove facts about the world, you will always want to apply judgment, skepticism, critical thinking, logical evaluation, intuition, etc. depending on the risk/reward tradeoff of the topic you're relying on the LLM for.
- t_mann 3y agoYeah I don't think it will ever make sense to think about Transformer models as 'understanding' something. The approach that I suggested would replace that with rather simple logic like answer_variance > arbitrary_threshold ? return 'I don't know' : return $original_answer It's not a fundamental fix, it doesn't even change the model itself, but the output might be more useful. And then there was just some speculation how you could try to train a new AI mimicking the more useful output. I'm sure smarter people than me can come up with way smarter approaches. But it wouldn't have to do with understanding - when I said the tool should return 'I don't know' above, I literally meant it should return that string (maybe augmented a bit by some pre-defined prompt), like a meaningless symbol, not any result of anything resembling introspection.
- gorlilla 3y agoYou left out hubris.
- andsoitis 3y ago> You left out hubris. I know!
- williamcotton 3y agoWe are having a conversation the feels much like the existence of a deity.
- andsoitis 3y ago> We are having a conversation the feels much like the existence of a deity. From a certain perspective, there does appear to be a rational mystical dualism at work.