3 ms·
That’s why production RAG systems really require full blown search backends. You decompose the natural language question into a search query with meta data filt
by lukebuehler 3y ago
That’s why production RAG systems really require full blown search backends. You decompose the natural language question into a search query with meta data filtering and more. And only then do embeddings search. And only if you get some good search results, with high confidence scores, you use an LLM to summarize or glue the results together, otherwise you branch and tell the user, “I don’t know”. IMO that’s the only way to make it work right now.