4 ms·
Thanks for the pointer, yeah the website content has gone stale. I'll try update it by end of day Khoj is using the Llama 7B, 4bit quantized, GGML by TheBloke.
by 110 3y ago
Thanks for the pointer, yeah the website content has gone stale. I'll try update it by end of day
Khoj is using the Llama 7B, 4bit quantized, GGML by TheBloke.
It's actually the first offline chat model that gives coherent answers to user queries given notes as context.
And it's interestingly more conversational than GPT3.5+, which is much more formal
- moneywoes 3y agoIs a vector db used?
- M4v3R 3y agoWhy only the 7B version? Would there be possibility to add support for the 13B as well if someone has enough RAM to run it?
- 110 3y agoWe're still trying to figure out the right balance between configurability and ease of use/maintenance. The 7B version was a decent enough starting point in terms of what it can answer (and way fewer folks can run a 13B on their machine). If you really want you can just replace the 7B model file with the 13B one under the ~/.cache/gpt4all directory on your device and it should just work.
- agg23 3y agoOh interesting, so you're not using Llama 2, you're using the original. Have you begun to evaluate Llama 2 to determine the differences in performance? How are you determining what notes (or snippets of notes?) to be injected as context? Especially given the small 2048 context limit with Llama 1.
- sabaimran 3y agoQuick clarification, we are using LlamaV2 7B. We didn't experiment with Llama 1 because we weren't sure of the licensing limitations. We determine note relevance by using cosine similarity between the query and the knowledge base (your note embeddings). We limit the context window for Llama2 to 3 notes (while OpenAI might comfortably take up to 9). The notes are ranked based on most to least similar and truncated based on the context window limit. For the model we're using, we're still limited to 2048 tokens for Llama v2.
- bugglebeetle 3y agoHave you looked at using the long context (32K) version of the Llama v2 7B released by Together AI? https://together.ai/blog/llama-2-7b-32k https://together.ai/blog/llama-2-7b-32k
- 110 3y agoOh neat, thanks for sharing that! Having a 32K offline model is pretty promising. Let me test out how it performs
- OkGoDoIt 3y agoI thought llama V2 has a context window of 4096?
- jmorgan 3y agoThis is a super cool project. Congrats! If you’re looking at trying different models with one API check out an open-source project a few folks and I have been working on in July in case it’s helpful https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama Llama 2 gives great answers, even the 7B model. There’s an “uncensored” 7B version as well George Sung has fine-tuned for topics that the default Llama2 model won’t discuss - eg I had trouble having Llama2 review authentication/security code or topics: https://huggingface.co/TheBloke/llama2_7b_chat_uncensored-GGML https://huggingface.co/TheBloke/llama2_7b_chat_uncensored-GG... From just playing around with it the uncensored model still seems to know where to “draw the line” on sensitive topics but YMMV If you do end up checking out Ollama you can try it with with this command or there’s an API too (it’s not in the docs yet) ollama run llama2-uncensored
- sabaimran 3y agoThis (ollama) is neat! Thanks for the pointers. Yeah, I ran into a couple of funny edge cases using Llama v2 with my personal notes. For example, if I ever asked it anything remotely personal (as I would with a personal assistant), it would often start telling me that asking for personal data is unethical. I get it, you have to be careful with the open source LLMs, but still a bit funny. It does work with enough coaxing though.
- xcdzvyn 3y agoThat's actually pretty embarrassing given the product's purpose. I believe there are uncensored LLAMA models out there that might be worth a shot