3 ms·
This is just "citing the sources" of the retrieved documents or am I missing something? I was expecting to read some novel approach like Hypothetical Document E
by mfalcon 3y ago
This is just "citing the sources" of the retrieved documents or am I missing something? I was expecting to read some novel approach like Hypothetical Document Embeddings.
- crdrost 3y agoYeah it's not a novel approach, it's 1. Chunkify the knowledge database into smallish chunks, 2. When you get a question, find the chunks most similar to the question 3. Prompt the LLM to use [the found chunks] to answer [the question] If you've ever posted a question to Stack Overflow or a similar site, think about the noise reduction “these posts may be relevant to you” feature. When you're actually asking an interesting and/or open ended question, you get a looooot of false positives. That's in the nature of text similarity search, it really hinges on strong nouns and verbs to anchor the discussion. It's no coincidence that the examples given here are uncommon names like Mussolini and Ketanji. So while this is not terribly useful, it is interesting that it might severely reduce the hallucination rate, which is a sort of false positive rate. Does it only reduce the false positives for facts that it “knows” via the external database? The really interesting prompts and responses, where you ask it to answer questions about figures who do not exist or data too npew for it to have, see if it still hallucinates, would have been very nice to see in this blog post. Missed opportunity.