4 ms·
I had several issues with unsloth gguf, even for models released a few months back like gemma 4, I have 0 confidence in their models, at this stage, I feel seve
by Foobar8568 2mo ago
I had several issues with unsloth gguf, even for models released a few months back like gemma 4, I have 0 confidence in their models, at this stage, I feel several uncensored are more reliable.
- segmondy 2mo ago... because they are often the first to quant it. sometimes the actually model providers will release wrong chat templates or values in the model config which leads to bad quants. how would you know a quant is good if you don't make one? you don't. so they make it first, then they run a lot of tests, KD, perplexity, etc, they publish it. They take feedback from the community, then they update if needed. if you want to try it right now, you grab it else wait for a week or 2.
- Foobar8568 2mo agoGemma4 was released a few weeks ago? The problems are still there. Today I started using another "provider" and the problems disappeared. Thanks but no.
- arcanemachiner 2mo agoThey are very responsive, and would probably be happy to help you fix your issues.
- danielhanchen 2mo agoHey yes - if you could describe what the issues are - we will gladly fix them!
- Foobar8568 2mo agoMea-culpa, dry-multiplier generated crap and even more so on Gemma4.
- danielhanchen 2mo agoOk no worries - if there are any future issues - feel free to message / make a HF issue - we'll fix promptly! Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes works, but it actually screws up reasoning sometimes
- segmondy 2mo agoI've 0 issues with gemma4 and I downloaded it early.
- LeBit 2mo agoSame. Used the 12B, 26B A4B and 31B. No issues with llama.cpp.
- jokethrowaway 2mo agoI've used their gemma 4 quants since when they were not still working in llama.cpp and ik-llama.cpp and I don't remember any problems They are the most reliable in my experience, but if you have alternatives you trust I'd love to know
- danielhanchen 2mo agoHey sorry what are the problems that you're experiencing - we're more than happy to help fix them!
- LeBit 2mo agoThank you for what you are doing.
- danielhanchen 2mo agoThanks for the support and to the community!
- chlorion 2mo agoHuh I have had great luck with unsloth quants so far. What issues are you having?
- SwellJoe 2mo agoCounterpoint: I've been using the Unsloth Gemma 4 quants extensively since very soon after release (on ROCm and Apple Silicon, I don't have any Nvidia hardware big enough), pretty much every quantization down to 4 bits (the QAT is the business, indistinguishable from the full-fat version, runs great on a slightly chonky desktop or laptop), and I haven't had any issues. The reason I use Gemma 4 so much often comes down to how reliable it is; when I want to experiment with llama.cpp settings, MTP, n-gram, etc. it's my go-to because I know there isn't anything wrong with the model or the quantizations that could interfere with the experiment. It did take a little while for Unsloth to update the Laguna S 2.1 quants to fix the yarn_attn_factor, and so it was a bit frustrating getting that quantization running right, but almost always, I pick the unsloth quantization if there is one. (Still waiting/hoping for a Ling 3.0 Flash.)
- danielhanchen 2mo agoWe will investigate Ling!
- BonerWiener 2mo ago> I feel several uncensored are more reliable. I have had the same experience with gemma 4 on same tasks being refused. But this is when working with cyber offensive tasks and the like. It excels in coding and is very fast on consumer hardware. So I would say use the right tool for the right task. What uncensored models can you recommend?
- jacquesm 2mo agoWithout substantiating what issues you have with unsloth gguf files you are just adding noise, no signal.