4 ms·
It's possible to build RAG pipelines that support answering complex questions over multiple documents, and I don't think I would say the whole architecture need
by zacmps 3y ago
It's possible to build RAG pipelines that support answering complex questions over multiple documents, and I don't think I would say the whole architecture needs to change.
Doing it fast is another story, LLMs are pretty high latency at the moment.
- _pdp_ 3y agoIf you have any examples I can go through it will be greatly appreciated. My main concern is that all RAGs just look for the top N best matching results. That does not mean that the information is in there.
- lmeyerov 3y agoI'm not sure why all RAG implementations would do that? Ex: Louie.AI will use a combination of tools to answer a question (database queries, RAG vector index lookups, Python, interactive charts, ...), and will often do multiple attempts on the same tool until it decides it has exhausted the immediate use of that tool. Ex: Louie is even learning from usage, so as an analyst, if Louie stopped digging early and you decided to manually look further, Louie learns from this: It'll know in future sessions that it may be worth looking further in that kind of scenario. None of that, in isolation, is unique to Louie.AI: It's just part of what it means to do a 'full' agent implementation vs a langchain/llamaindex/openai wrapper. There's an interesting question around knowledge graph style questions here -- should an agent do iterative top-n vector similarity searches, or a single wider search over a knowledge graph, or maybe the documents should be combined into a knowledge graph and that's what's embedded. We're exploring a lot here in our bigger customer projects, and I can't say there's a clear universal answer...
- davedx 3y agoYup, it depends on the type of data being searched for and how the information is represented in the corpus. It also depends on how information is distributed: IME often longer legal documents have a lot of cross referencing, so you need multiple clauses to collate a single answer. It’s an interesting area.
- treprinum 3y agoIf you want better accuracy in the similarity search, you make RAG chunks smaller. You then need to calibrate similarity "thresholds" to know the probability distribution of relevant/correct/incorrect chunks for any given sentence and what the N should be to reach them. It's not going to be perfect but you can end up with 90% accuracy on average which tends to be better than many full-text search solutions. Moreover, querying a vector DB takes <0.5s and you can run the LLM in the streaming mode getting responses pretty quickly leading to the illusion of talking to a real human (especially if you also stream audio/video with it).