4 ms·
> This problem goes away with the retrieval augmented generation (RAG) model, which first performs a web search for relevant sources before producing a summary
by hackerlight 3y ago
> This problem goes away with the retrieval augmented generation (RAG) model, which first performs a web search for relevant sources before producing a summary of its findings. However, even in the GPT-4 RAG model, we find that up to 30% of statements made are not supported by any sources provided, with nearly half of responses containing at least one unsupported statement.
Could you do something like this. Chunk/separate the individual claims made by the LLM, and the associated sources from RAG. Then feed them into a second LLM one by one, asking "is this claim $X_i reflective of the source $Y_i?". Whatever this second LLM says, feedback the response to the original LLM and ask it to revise what it's saying if the answer is "No". Iterate until the second LLM says "Yes" to every separate claim. Not perfect but might help.
- simonw 3y agoGoogle Gemini actually has a feature like this, though it's hard to spot. There's a colorful G icon below each Gemini response - if you click it, a second process runs which attempts to "fact check" the claims from the original prompt response - it highlights them in different colors and adds citation links to them.
- abid786 3y agoLLMs are better at this and it’ll probably marginally improve the output quality but there can potentially be hallucinations (false positives or negatives) even in this evaluation task.