3 ms·
Unsloth already has a guide for their quants: https://unsloth.ai/docs/models/qwen3.8 https://unsloth.ai/docs/models/qwen3.8
by ZeroCool2u 2mo ago
Unsloth already has a guide for their quants: https://unsloth.ai/docs/models/qwen3.8 https://unsloth.ai/docs/models/qwen3.8
- codedokode 2mo agoI wonder who is unsloth and where they got time, hardware and knowledge to quantize them?
- NitpickLawyer 2mo agounshloth started as a finetuning library with lots of optimisations so you could finetune on lower end hardware. Kind of OGs of the local community. Started by two brothers Michael and Daniel(?) a math wiz and a community builder/communicator. They've since gotten some VC backing, are active in quantising lots of models on release day (work w/ labs to prepare things), known for their optimised quants (use different bits for different layers). Recently I saw they launched some sort of a desktop app, like lmstudio if you're familiar with it. They're really cool people and known in the local model places.
- ZeroCool2u 2mo agoIt's our guy Daniel: https://www.linkedin.com/in/danielhanchen https://www.linkedin.com/in/danielhanchen
- danielhanchen 2mo agoHey :)
- Eisenstein 2mo agoThey started with offering training methods for quantized models to save memory and added new things over time. They are very active in the local model community and have extensive documentation and tooling to help with running and training models locally.
- arthurcolle 2mo agoDaniel Han is just that good!
- danielhanchen 2mo agoThanks haha
- kittikitti 2mo agoI quantize my models with llama.cpp and it's usually one command. Some of their quants are fine-tuned by architecture but it's only to squeeze out every little performance benefit.
- suprjami 2mo agoUnsloth imatrix data puts their quants at lower KLD than almost all others. It's true they make architecture-specific changes like keeping certain layers at F16 but it's also more than that.
- Foobar8568 2mo agoI had several issues with unsloth gguf, even for models released a few months back like gemma 4, I have 0 confidence in their models, at this stage, I feel several uncensored are more reliable.
- segmondy 2mo ago... because they are often the first to quant it. sometimes the actually model providers will release wrong chat templates or values in the model config which leads to bad quants. how would you know a quant is good if you don't make one? you don't. so they make it first, then they run a lot of tests, KD, perplexity, etc, they publish it. They take feedback from the community, then they update if needed. if you want to try it right now, you grab it else wait for a week or 2.
- Foobar8568 2mo agoGemma4 was released a few weeks ago? The problems are still there. Today I started using another "provider" and the problems disappeared. Thanks but no.
- arcanemachiner 2mo agoThey are very responsive, and would probably be happy to help you fix your issues.
- danielhanchen 2mo agoHey yes - if you could describe what the issues are - we will gladly fix them!
- Foobar8568 2mo agoMea-culpa, dry-multiplier generated crap and even more so on Gemma4.
- danielhanchen 2mo ago
- unleaded 2mo agoWhat hardware would you even be able to run this on?
- adrian_b 2mo agoEven the lowliest hardware could run this, but at an unlikely to be useful low speed, e.g. of 3 or 4 tokens per minute (by reading the weights from a couple of 4 TB SSDs for the BF16 model, or from a 4 TB SSD for the FP8 variant). The question about LLMs is never whether they can be run, because that has a trivial answer, they can always be run. The right question is what speeds are achievable for representative hardware configurations. At launch, it is difficult to estimate the speed. That should be known after someone reports experimental results. Moreover, for many LLMs the speed improved sometimes later after their release, after tweaks in inference backends, like llama.cpp or vLLM.
- badcafe23423435 2mo agotell me how run 35B on my 8GiB VRAM (linux) speed is not problem when You run agents and forget for 2-3 days