3 ms·
The tutorials, examples, demos, and blogs for every vector DB or RAG system that I know of focus mostly on toy datasets far, far under a million tokens. So most
by analyte123 3y ago
The tutorials, examples, demos, and blogs for every vector DB or RAG system that I know of focus mostly on toy datasets far, far under a million tokens. So most people can be forgiven for the view that long context would kill RAG - and it really should for these trivial use cases.
I don't know if vector DB and RAG vendors don't demo their software with legitimately large datasets because they don't want to compete with their customers, or because they're not confident enough in the results, or because they don't really use their software themselves.
To give an example of a RAG dataset I have played with, it is about 10k documents and 5M tokens. Adding prior versions, expanding the coverage, or augmenting from other sources I'm sure it could blow well past 10M tokens. Maybe extremely long context models will get extremely fast and cheap, but at least for the next couple years you probably won't be stuffing this all in the context. Even for a 1M context you still have to filter, retrieve, and rank the documents. And 10,000 documents is really not a lot compared to other corpora you could imagine (e-mails, media, law, science, code, etc). But it is certainly bigger than many of the use cases being pitched for RAG - like a few hundred personal notes or a corporate wiki with 500 poorly maintained pages on it.
- singularity2001 3y agoEven for small use cases with less than a million tokens RAG is often the better approach because of the time and money huge context windows consume: A minute (!!) and a dollar (?) per query in the Gemini demo