4 ms·
RAG approaches should work quite well for the examples you mentioned. It's a matter of how you approach the retrieval part - you can opt for a larger recall on
by armcat 3y ago
RAG approaches should work quite well for the examples you mentioned. It's a matter of how you approach the retrieval part - you can opt for a larger recall on retrieval, and leverage the large context window for the LLM to figure out the answer. Even if it's not "as-is", semantically if it's in there, it should be able to find it.
Other things to try out is how you approach the question "expansion" part, for example using Hypothetical Document Embeddings (HyDE); or how you approach the filtering-out part, e.g. using "System 2 Attention", https://arxiv.org/abs/2311.11829 https://arxiv.org/abs/2311.11829.
- curl-up 3y agoI tried most of such techniques, but the point is that this information really isn't in there directly, and to perform the question expansion, the model needs to know about the domain already. For example, imagine that one paper is about how author X was French, in early 19th-c.and how they were one of the first ones to write about topic T. Another paper is about how author Y was inspired by the early 19th-c. French writers writing about T. However, this second article does not mention X at all. Asking about "who were the main influences on X" would not give you the second article. Of course, I could run "multiple-hop" RAG-like process, where the model keeps asking questions itself and so on in a loop, but this becomes extremely clumsy, and the models (even GPT-4) tend to get out of hand. It is also extremely slow, of course.
- armcat 3y agoI've worked with EXACT this type of problem and for me RAG works perfectly well - it may seem "clumsy" as you put it in terms of trying to engineer or optimize the indexing, augmentation, and retrieval techniques, but it's worth it. I would claim that going down the finetuning route lends to a significantly higher probability of going astray at a MUCH larger cost. There have been some attempts to compare RAG vs FT approaches, I would recommend this paper: https://arxiv.org/abs/2403.01432 https://arxiv.org/abs/2403.01432
- curl-up 3y agoThanks for the input! Did you implement it in this kind of "multi-hop" way, or is there some trick I'm missing?