2 ms·
That assumes one layer of memory. In my experience you need to have at least 4 layers of memory to work well. All of them have different requirements for retrie
by jsemrau 5mo ago
That assumes one layer of memory. In my experience you need to have at least 4 layers of memory to work well. All of them have different requirements for retrieval.
Everything that is in short-term memory (state of the app, current conversation, current workspace artefact) requires fast latency and precision. For example if you want to edit a segment in a financial analysis, a blog post, or a program you only want to edit this segment. RAG on a VectorDB is overkill in my opinion.