6 ms·
Quick clarification, we are using LlamaV2 7B. We didn't experiment with Llama 1 because we weren't sure of the licensing limitations. We determine note relevan
by sabaimran 3y ago
Quick clarification, we are using LlamaV2 7B. We didn't experiment with Llama 1 because we weren't sure of the licensing limitations.
We determine note relevance by using cosine similarity between the query and the knowledge base (your note embeddings). We limit the context window for Llama2 to 3 notes (while OpenAI might comfortably take up to 9). The notes are ranked based on most to least similar and truncated based on the context window limit. For the model we're using, we're still limited to 2048 tokens for Llama v2.
- bugglebeetle 3y agoHave you looked at using the long context (32K) version of the Llama v2 7B released by Together AI? https://together.ai/blog/llama-2-7b-32k https://together.ai/blog/llama-2-7b-32k
- 110 3y agoOh neat, thanks for sharing that! Having a 32K offline model is pretty promising. Let me test out how it performs
- OkGoDoIt 3y agoI thought llama V2 has a context window of 4096?