3 ms·
A remaining advantage of large language models is that as they get larger, they tend to hallucinate less, simply because the odds of the training set containing
by Animats 2mo ago
A remaining advantage of large language models is that as they get larger, they tend to hallucinate less, simply because the odds of the training set containing a desired answer improve with size. If a solid "I don't know" detector is developed for inference, then you can try a small language model first.
An implication is that successful research in "I don't know" detection could destroy hundreds of billions in shareholder value.
- embedding-shape 2mo agoAnother "cool but we don't know how yet" thing would be a "confidence interval" so we know how much to trust LLM responses. Or while we're fantasizing, they could just know everything all the time regardless of training data. The "if a solid" part is easy to imagine, hard to implement :)
- root-parent 2mo ago>> A remaining advantage of large language models is that as they get larger, they tend to hallucinate less First time I hear that...not really true. "Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors" - https://arxiv.org/abs/2607.00447 https://arxiv.org/abs/2607.00447 "Calibrated Language Models Must Hallucinate" - https://arxiv.org/abs/2311.14648 https://arxiv.org/abs/2311.14648 "TruthfulQA: Measuring How Models Mimic Human Falsehoods" - https://arxiv.org/abs/2109.07958 https://arxiv.org/abs/2109.07958
- Animats 2mo agoFrom the "must hallucinate" paper: "For "arbitrary" facts whose veracity cannot be determined from the training data, we show that hallucinations must occur at a certain rate for language models that satisfy a statistical calibration condition appropriate for generative language models." The bigger the model, the more likely it is that arbitrary facts in the training data are embedded in the model. Then a larger model shows less hallucination on the same questions, since it has a matching answer stored for more questions. From the "TruthfulQA" paper: "Models generated many false answers that mimic popular misconceptions and have the potential to deceive humans. The largest models were generally the least truthful. This contrasts with other NLP tasks, where performance improves with model size. However, this result is expected if false answers are learned from the training distribution." That's more of a garbage-in, garbage out problem. If the large model is trained by shoveling in random web content, that's going to happen. Not a hallucination problem. The LLM just fed back what it had been told.