4 ms·
One caveat about about embedding based retrieval is that there is no guarantee that the embedded documents will look like the query. One trick is to have a LLM
by orasis 3y ago
One caveat about about embedding based retrieval is that there is no guarantee that the embedded documents will look like the query.
One trick is to have a LLM hallucinate a document based on the query, and then embed that hallucinated document. Unfortunately this increases the latency since it incurs another round trip to the LLM.
- williamcotton 3y ago“We’re gonna need a bigger boat.”
- rco8786 3y ago> One trick is to have a LLM hallucinate a document based on the query I'm not following why you would want to do this? At that point, just asking the LLM without any additional context would/should produce the same (inaccurate) results.
- BoorishBears 3y agoYou're not having the LLM answer from the hallucination, you're looking for the document that looks most similar to the hallucination and having it answer on that instead.
- taberiand 3y agoIs that something easily handed off to a faster/cheaper LLM? I'm imagining something like running the main process through GPT-4 and hand of the hallucinations to GPT 3 turbo. If you could spot the need for it while streaming a response you could possibly even have it ready ahead of time
- wasabi991011 3y ago>One caveat about about embedding based retrieval is that there is no guarantee that the embedded documents will look like the query. Aleph Alpha provides an asymmetric embedding model which I believe is an attempt to resolve this issue (haven't looked into it much, just saw the entry in langchain's documentation)
- redskyluan 3y agoi have an opposite way on doing this. Tried to generate questions based on doc chunks and embedding on questions. It works perfect!
- ck_one 3y agoHow do you generate the questions and how do you make sure to not lose information? E.g. Today I woke up at 9.am, had a light breakfast and then went on a run in Golden Gate Park. What questions do you generate from this sentence?
- redskyluan 3y agoThey generate questions like: where did you go this morning? When did you woke up this morning. What did you do after breakfast? What did you do today at Golden Gate Park. GPT is all about probabilities. So the LLM know what might be most related answer of a doc chunk. It works much better than embedding the whole sentence because "When did you woke up this morning" might not be very similar with "Today I woke up at 9.am, had a light breakfast and then went on a run in Golden Gate Park.".
- ck_one 3y agoInteresting, thank you! Have you played around with different prompts to generate these questions?
- redskyluan 3y agoThat's exactly what I'm trying to do. Play with prompts, generate multiple questions and cluster questions and pick some of the centroid questions to do embeddings
- orasis 3y agoNice! Do you generate N questions so N embeddings per document or just one?
- 3y ago
- d4rkp4ttern 3y agoSome people packaged this rather intuitive idea, named it Hyde (Hypothetical Document Embeddings) and wrote a paper about it — https://arxiv.org/abs/2212.10496 https://arxiv.org/abs/2212.10496 Summary — HyDE is a new method for creating effective zero-shot dense retrieval systems that generates hypothetical documents based on queries and encodes them using an unsupervised contrastively learned encoder to identify relevant documents. It outperforms state-of-the-art unsupervised dense retrievers and performs strongly compared to fine-tuned retrievers across various tasks and languages.