4 ms·
The described technique of RAG is not only expensive, but also prone to hallucinations. I would have liked more discussion on hallucinations, which is the ulti
by MattDaEskimo 3y ago
The described technique of RAG is not only expensive, but also prone to hallucinations.
I would have liked more discussion on hallucinations, which is the ultimate pitfall of LLMs. This is critical for discovery-based public-facing chatbots.
I'm also very skeptical of real-world HyDE applications as they depend on the underlying model to properly answer the question, and can easily drift from the intention.
- Der_Einzige 3y agoFat citation needed on RAG being expensive. Most embeddings models these days are smaller than most LLM models, and run more cheaply and more effectively than the ones provided by OpenAI. If you mean that increasing token counts in expensive, I suppose sure - but the retrieval side itself is not the cost center in general.
- MattDaEskimo 3y agoI am speaking of the OPs implementation of RAG, not in general. The retrieval part can be expensive if an LLM is used to confirm that it is sufficient, and if it's decided to continue searching if it's not. OpenAIs retrieval is a perfect example. It works, but it's very expensive
- simonw 3y agoI've seen a lot of people (including people who I trust) say RAG is the best current mitigation we have for hallucinations, because grounding the LLM in additional context makes it much less likely it will make something up as opposed to use the information that has been passed to it along with the user's question. Have you heard differently?
- simonw 3y agoRelated: just saw this paper https://arxiv.org/abs/2401.01313 https://arxiv.org/abs/2401.01313 "A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models"
- MattDaEskimo 3y agoRAG conceptually is the solution for hallucinations. I'm being critical of the implementations used to achieve it and the lack of awareness for hallucinations. It's definitely not ready.... Yet.
- zmmmmm 3y agoI spent a bit of time playing with h2ogpt, which is a popular RAG framework. I gave it all our architectural documentation for our software and then tried asking it questions that transcended basic search (so the answer is not directly in there, but you could formulate it if you put disparate parts together). It started hallucinating pretty fast and told me all kinds of BS about how our software works. I think RAG can be used in a way that eliminates or drastically reduces hallucinations, but to do that you have to do quite a lot of work to constrain the context and structure the prompting to address very specific questions. When you apply these more general frameworks they pump in large amounts of context in an unstructured manner and you just end right back at hallucinations again because the context isn't constrained enough. So RAG is useful to me but not a silver bullet. It doesn't solve the original problem of wanting all the features of an LLM but without the hallucinations. It gives you some targeted way to use the LLM that doesn't hallucinate but misses a lot of the functionality people want.