4 ms·
The models that exist now say "I don't know" all the time. It's so weird that people keep insisting that it can't do things that it does. Ask it what dark mat
by empath-nirvana 3y ago
The models that exist now say "I don't know" all the time. It's so weird that people keep insisting that it can't do things that it does.
Ask it what dark matter is, and it won't invent an answer, it will present existing theories and say that it's unknown.
Ask it about a person you know that isn't in it's data set and it'll tell you it has no information about the person.
Despite the fact that people insist that hallucinations are common and that it will invent answers if it doesn't know something frequently, the truth is that chatgpt doesn't hallucinate that much and will frequently say it doesn't know things.
One of the few cases where I've noticed it inventing things are that it often makes up apis for programming libraries and CLI tools that don't exist, and that's trivially fixable by referring it to documentation.
- intended 3y agoI have to use LLMs for work projects - which are not PoCs. I can’t have a tool that makes up stuff an unknown amount of time. There is a world of research examining hallucination Rates, indicating hallucination rates of 30%+. With steps to reduce it using RAGs, you could potentially improve the results significantly - last I checked it was 80-90%. And the failure types aren’t just accuracy, it’s precision, recall, relevance and more.
- empath-nirvana 3y ago> There is a world of research examining hallucination Rates, indicating hallucination rates of 30%+. I want to see a citation for this. And a clear definition for what is a hallucination and what isn't.
- intended 3y agohttps://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=llm+hallucination&oq=llm+hallucinatioN https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=llm+... https://www.medpagetoday.com/ophthalmology/generalophthalmology/105672#:~:text=The%20mean%20hallucination%20rate%20for,and%2029%25%20with%20version%204.0 https://www.medpagetoday.com/ophthalmology/generalophthalmol.... - Survey of Hallucination in Natural Language Generation](https://arxiv.org/abs/2202.03629 https://arxiv.org/abs/2202.03629) (Ji et al., 2022) - [How Language Model Hallucinations Can Snowball](https://arxiv.org/abs/2305.13534 https://arxiv.org/abs/2305.13534) (Zhang et al., 2023) - [A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity](https://arxiv.org/abs/2302.04023 https://arxiv.org/abs/2302.04023) (Bang et al., 2023) - [Contrastive Learning Reduces Hallucination in Conversations](https://arxiv.org/abs/2212.10400 https://arxiv.org/abs/2212.10400) (Sun et al., 2022) - [Self-Consistency Improves Chain of Thought Reasoning in Language Models](https://arxiv.org/abs/2203.11171 https://arxiv.org/abs/2203.11171) (Wang et al., 2022) - [SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models](https://arxiv.org/abs/2303.08896 https://arxiv.org/abs/2303.08896) ( Manakul et al., 2023)