4 ms·
Are you asking about the query/tokenization or the overall architecture? I assume it's the former. It's explained in the post after the diagram, see "The diagr
by nkaretnikov 3y ago
Are you asking about the query/tokenization or the overall architecture?
I assume it's the former. It's explained in the post after the diagram, see "The diagram illustrates a series of steps [...]".
We get a user query. Based on that query, we pull relevant parts from the doc. Then, we submit both the query and the doc parts to an LLM. This limits the amount of data we need to send and allows the user to know which parts of the doc are relevant.