4 ms·
Sorry to repeat the same question I just asked the other commenter in this thread, but could you link the model page and recommend a specific level of quantizat
by peter_l_downs 3y ago
Sorry to repeat the same question I just asked the other commenter in this thread, but could you link the model page and recommend a specific level of quantization for the models you've referenced? I'd love to play with these models and see what you're talking about.
- BOOSTERHIDROGEN 3y agoIt's from nous research https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B Q5 is minimum.
- peter_l_downs 3y agoThank you — from that page, at the bottom, I was able to find this link to what I think are the quantized versions https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B-GGUF/tree/main https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B-... If you have the time, could you explain what you mean by "Q5 is minimum"? Did you determine that by trying the different models and finding this one is best, or did someone else do that evaluation, or is that just generally accepted knowledge? Sorry, I find this whole ecosystem quite confusing still, but I'm very new and that's not your problem.
- BOOSTERHIDROGEN 3y agoIt's the best balance if you have limited compute performance.
- peter_l_downs 3y agoThank you
- d-z-m 3y agoTalking GGUF, Usually the higher you can afford to go wrt. quantization(e.g. Q5 is better than Q4, etc), the better. A Q6_K has minimal performance loss from the Q8, so in most cases if you can fit a Q6_K it's recommended to just use that. TheBloke's READMEs[0] usually have a good table summarizing each quantization level. If you're RAM constrained, you'll also have to make trade-offs about the context length. e.g. you could have 8 GB RAM and a Q5 quant with shorter context, vs Q3 with longer, etc. [0]:https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF
- peter_l_downs 3y agoThank you!