5 ms·
It depends what problem you're solving. If it's a high frequency request, like a chat response, it's far too inefficient. Most web APIs would consider it bad pr
by tkellogg 3y ago
It depends what problem you're solving. If it's a high frequency request, like a chat response, it's far too inefficient. Most web APIs would consider it bad practice to read 2MB of data on every request, even worse when you consider all the LLM computation. Instead, use RAG and pull targeted info out of some sort of low-latency database.
However, caching might be a sweet spot for these multi-modal and large context LLMs. Take a bunch of documents and perform reasoning tasks to distill the knowledge down into something like a knowledge graph, to be used in RAG.