4 ms·
Thank you — from that page, at the bottom, I was able to find this link to what I think are the quantized versions https://huggingface.co/NousResearch/Hermes-2
by peter_l_downs 3y ago
Thank you — from that page, at the bottom, I was able to find this link to what I think are the quantized versions
https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B-GGUF/tree/main https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B-...
If you have the time, could you explain what you mean by "Q5 is minimum"? Did you determine that by trying the different models and finding this one is best, or did someone else do that evaluation, or is that just generally accepted knowledge? Sorry, I find this whole ecosystem quite confusing still, but I'm very new and that's not your problem.
- BOOSTERHIDROGEN 3y agoIt's the best balance if you have limited compute performance.
- peter_l_downs 3y agoThank you
- d-z-m 3y agoTalking GGUF, Usually the higher you can afford to go wrt. quantization(e.g. Q5 is better than Q4, etc), the better. A Q6_K has minimal performance loss from the Q8, so in most cases if you can fit a Q6_K it's recommended to just use that. TheBloke's READMEs[0] usually have a good table summarizing each quantization level. If you're RAM constrained, you'll also have to make trade-offs about the context length. e.g. you could have 8 GB RAM and a Q5 quant with shorter context, vs Q3 with longer, etc. [0]:https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF
- peter_l_downs 3y agoThank you!