5 ms·
Thinking / reasoning + multimodal + tool calling. We made some quants at https://huggingface.co/collections/unsloth/gemma-4 https://huggingface.co/collections/
by danielhanchen 6mo ago
Thinking / reasoning + multimodal + tool calling.
We made some quants at https://huggingface.co/collections/unsloth/gemma-4 https://huggingface.co/collections/unsloth/gemma-4 for folks to run them - they work really well!
Guide for those interested: https://unsloth.ai/docs/models/gemma-4 https://unsloth.ai/docs/models/gemma-4
Also note to use temperature = 1.0, top_p = 0.95, top_k = 64 and the EOS is "<turn|>". "<|channel>thought\n" is also used for the thinking trace!
- l2dy 6mo agoFYI, screenshot for the "Search and download Gemma 4" step on your guide is for qwen3.5, and when I searched for gemma-4 in Unsloth Studio it only shows Gemma 3 models.
- danielhanchen 6mo agoWe're still updating it haha! Sorry! It's been quite complex to support new models without breaking old ones
- smallerize 6mo agoSpeaking of which, do you think Step 3.5 Flash is going to happen or should I stop holding my breath?
- danielhanchen 6mo agoOh quants - haha I can re-investigate it - just totally forgot about them
- Imustaskforhelp 6mo agoDaniel, I know you might hear this a lot but I really appreciate a lot of what you have been doing at Unsloth and the way you handle your communication, whether within hackernews/reddit. I am not sure if someone might have asked this already to you, but I have a question (out of curiosity) as to which open source model you find best and also, which AI training team (Qwen/Gemini/Kimi/GLM) has cooperated the most with the Unsloth team and is friendly to work with from such perspective?
- danielhanchen 6mo agoThanks a lot for the support :) Tbh Gemma-4 haha - it's sooooo good!!! For teams - Google haha definitely hands down then Qwen, Meta haha through PyTorch and Llama and Mistral - tbh all labs are great!
- Imustaskforhelp 6mo agoNow you have gotten me a bit excited for Gemma-4, Definitely gonna see if I can run the unsloth quants of this on my mac air & thanks for responding to my comment :-)
- danielhanchen 6mo agoThanks! Have a super good day!!
- evilelectron 6mo agoDaniel, your work is changing the world. More power to you. I setup a pipeline for inference with OCR, full text search, embedding and summarization of land records dating back 1800s. All powered by the GGUF's you generate and llama.cpp. People are so excited that they can now search the records in multiple languages that a 1 minute wait to process the document seems nothing. Thank you!
- danielhanchen 6mo agoOh appreciate it! Oh nice! That sounds fantastic! I hope Gemma-4 will make it even better! The small ones 2B and 4B are shockingly good haha!
- qingcharles 6mo agoJust switched from 3.1 Flash Lite to Gemma-4 31B on the AI Studio API since there is a generous 1500/day on non-billed projects. It's doing fantastic.
- polishdude20 6mo agoHey in really interested in your pipeline techniques. I've got some pdfs I need to get processed but processing them in the cloud with big providers requires redaction. Wondering if a local model or a self hosted one would work just as well.
- jorl17 6mo agoSeconded, would also love to hear your story if you would be willing
- evilelectron 6mo agoI run llama.cpp with Qwen3-VL-8B-Instruct-Q4_K_S.gguf with mmproj-F16.gguf for OCR and translation. I also run llama.cpp with Qwen3-Embedding-0.6B-GGUF for embeddings. Drupal 11 with ai_provider_ollama and custom provider ai_provider_llama (heavily derived from ai_provider_ollama) with PostreSQL and pgvector. People on site scan the documents and upload them for archival. The directory monitor looks for new files in the archive directories and once a new file is available, it is uploaded to Drupal. Once a new content is created in Drupal, Drupal triggers the translation and embedding process through llama.cpp. Qwen3-VL-8B is also used for chat and RAG. Client is familiar with Drupal and CMS in general and wanted to stay in a similar environment. If you are starting new I would recommend looking at docling.
- zaat 6mo agoThank you for your work. You have an answer on your page regarding "Should I pick 26B-A4B or 31B?", but can you please clarify if, assuming 24GB vRAM, I should pick a full precision smaller model or 4 bit larger model?
- danielhanchen 6mo agoThank you! I presume 24B is somewhat faster since it's only 4B activated - 31B is quite a large dense model so more accurate!
- ryandrake 6mo agoThis is one of the more confusing aspects of experimenting with local models as a noob. Given my GPU, which model should I use, which quantization of that model should I pick (unsloth tends to offer over a dozen!) and what context size should I use? Overestimate any of these, and the model just won't load and you have to trial-and-error your way to finding a good combination. The red/yellow/green indicators on huggingface.co are kind of nice, but you only know for sure when you try to load the model and allocate context.
- danielhanchen 6mo agoDefinitely Unsloth Studio can help - we recommend specific quants (like Gemma-4) and also auto calculate the context length etc!
- ryandrake 6mo agoWill have to try it out. I always thought that was more for fine-tuning and less for inference.
- danielhanchen 6mo agoOh yes sadly we partially mis-communicated haha - there's both and synthetic data generation + exporting!
- pentagrama 6mo agoHey, I tried to use Unsloth to run Gemma 4 locally but got stuck during the setup on Windows 11. At some point it asked me to create a password, and right after that it threw an error. Here’s a screenshot: https://imgur.com/a/sCMmqht https://imgur.com/a/sCMmqht This happened after running the PowerShell setup, where it installed several things like NVIDIA components, VS Code, and Python. At the end, PowerShell tell me to open a http://localhost URL in my browser, and that’s where I was prompted to set the password before it failed. Also, I noticed that an Unsloth icon was added to my desktop, but when I click it, nothing happens. For context, I’m not a developer and I had never used PowerShell before. Some of the steps were a bit intimidating and I wasn’t fully sure what I was approving when clicking through. The overall experience felt a bit rough for my level. It would be great if this could be packaged as a simple .exe or a standalone app instead of going through terminal and browser steps. Are there any plans to make something like that?
- danielhanchen 6mo agoApologies we just fixed it!! If you try again from source ie irm https://unsloth.ai/install.ps1 https://unsloth.ai/install.ps1 | iex it should work hopefully. If not - please at us on Discord and we'll help you! The Network error is a bummer - we'll check. And yes we're working on a .exe!!
- pentagrama 6mo agoIt worked! https://imgur.com/a/SOfiRhv https://imgur.com/a/SOfiRhv Thanks, will check it out tomorrow. Hope the unsloth-setup.exe > Windows App is coming soon! I think it will expand accessibility and user base.
- danielhanchen 6mo agoOh nice! Glad it worked! Yes!! We're working on the app!
- deleted 6mo ago[deleted]
- nnucera 6mo agoWow! Thank you very much!
- danielhanchen 6mo agoThanks!
- egeres 6mo agoThank you and your brother for all the amazing work, it's really inspiring to others <3
- danielhanchen 6mo agoThank you and appreciate it!
- jquery 6mo agoAwesome!! Thank you SO much for this.
- danielhanchen 6mo agoAppreciate it!
- Wowfunhappy 6mo agoHi! Do you ever make quants of the base models? I'm interested in experimenting with them in non-chat contexts.
- car 6mo agoYes, they are listed on huggingface. The instruction trained models have an 'it' in their name. https://huggingface.co/collections/unsloth/gemma-4 https://huggingface.co/collections/unsloth/gemma-4 Edit: Sorry, I'm not sure if this is a quant, but it says 'finetuned' from the Google Gemma 4 parent snapshot. It's the same size as the UD 8-bit quant though.
- Wowfunhappy 6mo agoOnly the 'it' models seem to have quants. I was really hoping to try a base model.
- kristjansson 6mo agoBasic quantization is easy if you have enough RAM (not VRAM) to load the weights.
- zobzu 6mo agoneat, time to update my spam filter model hehe
- danielhanchen 6mo agoHaha! Ye the model is really good
- Kye 6mo agoI haven't tried a local model in a while. I can only fit E4B in VRAM (8GB), but it's good enough that I can see it replacing Claude.ai for some things.
- akavel 6mo agoI'm trying to disable "thinking", but it doesn't seem to work (in llama.cpp). The usual `--reasoning-budget 0` doesn't seem to change it, nor `--chat-template-kwargs '{"enable_thinking":false}'` (both with `--jinja`). Am I missing something? EDIT: Ok, looks like there's yet another new flag for that in llama.cpp, and this one seems to work in this case: `--reasoning off`. FWIW, I'm doing some initial tries of unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL, and for writing some Nix, I'm VERY impressed - seems significantly better than qwen3.5-35b-a3b for me for now. Example commandline on a Macbook Air M4 32gb RAM: llama-cli -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL -t 1.0 --top-p 0.95 --top-k 64 -fa on --no-mmproj --reasoning-budget 0 -c 32768 --jinja --reasoning off (at release b8638, compiled with Nix)
- danielhanchen 6mo agoOh very cool! Will check the `--reasoning off` flag as well! Yep the models are really good!
- kapimalos 6mo agoNoob question. Why I would use this version over the original model?
- piyh 6mo ago1/3 the RAM & CPU consumed for 99% the performance
- trashcan2137 6mo agoand the EOS is "<turn|>". "<|channel>thought\n" is also used for the thinking trace! Can someone explain this to me? Why is this faux-XML important here?
- pertymcpert 6mo agoThat’s how the model is trained to signal the end to its generation and to indicate its thinking.
- sroussey 6mo agoThese are likely individual tokens. They are super common.
- genpfault 6mo agollama.cpp (b8642) auto-fits ~200k context on this 24GB RX 7900 XTX & it shows a solid 100+ tok/s ("S_TG t/s") on the first 32k of it, nice! ./llama-batched-bench -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \ -npp 1000,2000,4000,8000,16000,32000,64000,96000,128000 -ntg 128 -npl 1 -c 0 | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s | |-------|--------|------|--------|----------|----------|----------|----------|----------|----------| | 1000 | 128 | 1 | 1128 | 0.416 | 2404.87 | 1.064 | 120.29 | 1.480 | 762.20 | | 2000 | 128 | 1 | 2128 | 0.755 | 2649.86 | 1.075 | 119.04 | 1.830 | 1162.83 | | 4000 | 128 | 1 | 4128 | 1.501 | 2665.72 | 1.093 | 117.08 | 2.594 | 1591.49 | | 8000 | 128 | 1 | 8128 | 3.142 | 2545.85 | 1.114 | 114.87 | 4.257 | 1909.47 | | 16000 | 128 | 1 | 16128 | 6.908 | 2316.00 | 1.189 | 107.65 | 8.097 | 1991.73 | | 32000 | 128 | 1 | 32128 | 16.382 | 1953.31 | 1.278 | 100.12 | 17.661 | 1819.16 | | 64000 | 128 | 1 | 64128 | 43.427 | 1473.74 | 1.453 | 88.12 | 44.879 | 1428.89 | | 96000 | 128 | 1 | 96128 | 82.227 | 1167.50 | 1.623 | 78.86 | 83.850 | 1146.42 | |128000 | 128 | 1 | 128128 | 133.237 | 960.69 | 1.797 | 71.25 | 135.034 | 948.86 |
- danielhanchen 6mo agoOh nice that's pretty good!
- spwa4 6mo ago~50 tok/s on M1 Max 64Gb
- sillysaurusx 6mo agoTemperature 1.0 used to be bad for sampling. 0.7 was the better choice, and the difference in results were noticeable. You may want to experiment with this.
- danielhanchen 6mo agoYou might be right, but Google's recommendation was temp 1 etc primarily because all their benchmarks were used with these numbers, so it's better reproducibility for downstream tasks
- sillysaurusx 6mo agoFair, though putting a note in the readme about temperature 0.7 couldn't hurt. I wonder why they do benchmarks with 1 instead of 0.7... that's strange. 0.7 or 0.8 at most gives noticeably better samples.
- davedx 6mo agoReproducibility. They're benchmarks.
- sillysaurusx 6mo agoReproducibility is a matter of using the same input seeds, which jax can do. 0.7 vs 1.0 would make no difference for that. Without seeds, 0.7 would be less random than 1.0, so it'd be (slightly) more reproducible.
- sixhobbits 6mo agoThanks for this, I gave this guide to my Claude and he oneshot the unsloth and gemma4 set up on the old macbook he runs on. It's way faster than I expected, haven't tried out local models for a few generations but will be very nice when they become useful
- danielhanchen 6mo agoThanks! Oh nice! Ye local models are advancing much faster than I expected!
- zkmon 6mo agoHow does Gemma 4 26B A4B compare with Qwen3.5 35B A3B for same quants(4)
- deleted 6mo ago[deleted]
- rizzo94 6mo agoHuge fan of the Unsloth quants! Having reasoning and tool calling this accessible locally is a massive leap forward. The main hurdle I've found with local tool calling is managing the execution boundaries safely. I’ve started plugging these local models into PAIO to handle that. Since it acts as a hardened execution layer with strict BYOK sovereignty, it lets you actually utilize Gemma-4's tool calling capabilities without the low-level anxiety of a hallucination accidentally wiping your drive. It’s the perfect secure gateway for these advanced local models.
- mmaunder 6mo agoThis comment deserves it's own HN post. Thanks!