4 ms·
Does anyone have sources, or experience, about fine tuning primarily to teach the model some factual data, especially when it comes to later "higher level" ques
by curl-up 3y ago
Does anyone have sources, or experience, about fine tuning primarily to teach the model some factual data, especially when it comes to later "higher level" question answering.
For example, giving the model a bunch of text (academic papers and such) about 19th century writers, then asking things like "Who were the main influences on writer X"?
Obviously simple RAG-like approaches don't work, as such information is rarely available in the text as-is, and needs to be "extrapolated" to some extent. Long context models might work (just dumping everything into the prompt), but are way too expensive for my needs.
- armcat 3y agoRAG approaches should work quite well for the examples you mentioned. It's a matter of how you approach the retrieval part - you can opt for a larger recall on retrieval, and leverage the large context window for the LLM to figure out the answer. Even if it's not "as-is", semantically if it's in there, it should be able to find it. Other things to try out is how you approach the question "expansion" part, for example using Hypothetical Document Embeddings (HyDE); or how you approach the filtering-out part, e.g. using "System 2 Attention", https://arxiv.org/abs/2311.11829 https://arxiv.org/abs/2311.11829.
- curl-up 3y agoI tried most of such techniques, but the point is that this information really isn't in there directly, and to perform the question expansion, the model needs to know about the domain already. For example, imagine that one paper is about how author X was French, in early 19th-c.and how they were one of the first ones to write about topic T. Another paper is about how author Y was inspired by the early 19th-c. French writers writing about T. However, this second article does not mention X at all. Asking about "who were the main influences on X" would not give you the second article. Of course, I could run "multiple-hop" RAG-like process, where the model keeps asking questions itself and so on in a loop, but this becomes extremely clumsy, and the models (even GPT-4) tend to get out of hand. It is also extremely slow, of course.
- armcat 3y agoI've worked with EXACT this type of problem and for me RAG works perfectly well - it may seem "clumsy" as you put it in terms of trying to engineer or optimize the indexing, augmentation, and retrieval techniques, but it's worth it. I would claim that going down the finetuning route lends to a significantly higher probability of going astray at a MUCH larger cost. There have been some attempts to compare RAG vs FT approaches, I would recommend this paper: https://arxiv.org/abs/2403.01432 https://arxiv.org/abs/2403.01432
- curl-up 3y agoThanks for the input! Did you implement it in this kind of "multi-hop" way, or is there some trick I'm missing?