3 ms·
> I feel like using LLM today is like using search 15 years ago - you get a feel for getting results you want. I don't think it's quite the same. With search
by TerrifiedMouse 3y ago
> I feel like using LLM today is like using search 15 years ago - you get a feel for getting results you want.
I don't think it's quite the same.
With search results, aka web sites, you can compare between them and get a "majority opinion" if you have doubts - it doesn't guarantee correctness but it does improve the odds.
Some sites are also more reputable and reliable than others - e.g. if the information is from Reuters, a university's courseware, official government agencies, ... etc. it's probably correct.
With LLMs you get one answer and that's it - although some like Bard provide alternate drafts but they are all from the same source and can all be hallucinations ...
- famouswaffles 3y ago>although some like Bard provide alternate drafts but they are all from the same source and can all be hallucinations ... Yes and no. If the LLM is repeating the same thing on multiple drafts then it's very unlikely to be a hallucination. It's when multiple generations are all saying different things that you need to take notice. LLMs hallucinate yes but getting the same hallucination multiple times is incredibly rare.
- PoignardAzur 3y agoWait, is that true? I feel like that claim needs a lot of disclaimers.
- famouswaffles 3y agohttps://arxiv.org/abs/2305.18248 https://arxiv.org/abs/2305.18248 "In particular, we find that LMs often hallucinate differing authors of hallucinated references when queried in independent sessions, while consistently identify authors of real references. This suggests that the hallucination may be more a generation issue than inherent to current training techniques or representation." https://arxiv.org/abs/2303.08896 https://arxiv.org/abs/2303.08896 "SelfCheckGPT leverages the simple idea that if a LLM has knowledge of a given concept, sampled responses are likely to be similar and contain consistent facts. However, for hallucinated facts, stochastically sampled responses are likely to diverge and contradict one another."
- TerrifiedMouse 3y agoThen why aren’t hallucinations being eliminated by comparing drafts?
- famouswaffles 3y agoautomatically comparing drafts for every single query would be expensive. and that wouldn't eliminate hallucinations just tell you if large details have likely been hallucinated. But it's a method some research has used. https://arxiv.org/abs/2303.08896 https://arxiv.org/abs/2303.08896
- TerrifiedMouse 3y agoHow expensive could it be? Google Bard, a free service, offers the drafts for free. Just do the comparison on the user’s machine if the LLM provider is that cheap. P.S. Also aren’t LLMs deterministic if you set their “temperature” to zero? Are there drafts if the temperature is zero? If not, then that’s the same as removing the randomness no?
- famouswaffles 3y agoThe drafts have to be evaluated either by a human or llm. Doing that for every request does not scale when you have millions of users. >Just do the comparison on the user’s machine if the LLM provider is that cheap. This is not possible. Users don't have the resources to run these gigantic models. LLM inference is not cheap. Open ai, Google aren't running profit on free cGPT or Bard. >P.S. Also aren’t LLMs deterministic if you set their “temperature” to zero? Are there drafts if the temperature is zero? If not, then that’s the same as removing the randomness no? It's not a problem of randomness. a temp of 0 doesn't reduce hallucinations. LLMs internally know when they are hallucinating/taking a wild guess. randomness influences how that guess manifests each time but the decision to guess was already made. https://arxiv.org/abs/2304.13734 https://arxiv.org/abs/2304.13734
- TerrifiedMouse 3y ago