3 ms·
100% I think the author is really misunderstanding the issue here. "Hallucination" is a fundamental aspect of the design of Large Language Models. Narrowing the
by dweinus 2y ago
100% I think the author is really misunderstanding the issue here. "Hallucination" is a fundamental aspect of the design of Large Language Models. Narrowing the distribution of the training data will reduce the LLM's ability to generalize, but it won't stop hallucinations.
- sean_pedersen 2y agoI agree in that a perfectly consistent dataset won't completely stop statistical language models from hallucinating but it will reduce it. I think it is established that data quality is more important than quantity. Bullshit in -> bullshit out, so a focus on data quality is good and needed IMO. I am also saying LMs output should cite sources and give confidence scores (which reflects how much the output is in or out of the training distrtibution).
- rtkwe 2y agoI think the problem is you need an extremely large quantity of data just to get the machine to work in the first place. So much so that there may not be enough to get it working on just "quality" data.
- bbor 2y agoWhat’s a non-statistical language model? And I think looking to the training data for sources is a little silly - that’s the training data for intuitive language use, not true statements about the world. If you haven’t checked it out yet, two terms you’d love are “RAG” and “Manuel De Landa”
- antihipocrat 2y agoHow would confidence scores work? Multiple passthroughs and a % attached to each statement according to how often it appeared in the generated result? If so, building this could be quite complex depending on the domain. In the legal field even one simple word that is changed can have large consequences.