12 ms·
Google releases Gemma 4 open models
- danielhanchen 6mo agoThinking / reasoning + multimodal + tool calling. We made some quants at https://huggingface.co/collections/unsloth/gemma-4 https://huggingface.co/collections/unsloth/gemma-4 for folks to run them - they work really well! Guide for those interested: https://unsloth.ai/docs/models/gemma-4 https://unsloth.ai/docs/models/gemma-4 Also note to use temperature = 1.0, top_p = 0.95, top_k = 64 and the EOS is "<turn|>". "<|channel>thought\n" is also used for the thinking trace!
- l2dy 6mo agoFYI, screenshot for the "Search and download Gemma 4" step on your guide is for qwen3.5, and when I searched for gemma-4 in Unsloth Studio it only shows Gemma 3 models.
- danielhanchen 6mo agoWe're still updating it haha! Sorry! It's been quite complex to support new models without breaking old ones
- smallerize 6mo agoSpeaking of which, do you think Step 3.5 Flash is going to happen or should I stop holding my breath?
- danielhanchen 6mo agoOh quants - haha I can re-investigate it - just totally forgot about them
- Imustaskforhelp 6mo agoDaniel, I know you might hear this a lot but I really appreciate a lot of what you have been doing at Unsloth and the way you handle your communication, whether within hackernews/reddit. I am not sure if someone might have asked this already to you, but I have a question (out of curiosity) as to which open source model you find best and also, which AI training team (Qwen/Gemini/Kimi/GLM) has cooperated the most with the Unsloth team and is friendly to work with from such perspective?
- danielhanchen 6mo agoThanks a lot for the support :) Tbh Gemma-4 haha - it's sooooo good!!! For teams - Google haha definitely hands down then Qwen, Meta haha through PyTorch and Llama and Mistral - tbh all labs are great!
- Imustaskforhelp 6mo agoNow you have gotten me a bit excited for Gemma-4, Definitely gonna see if I can run the unsloth quants of this on my mac air & thanks for responding to my comment :-)
- danielhanchen 6mo agoThanks! Have a super good day!!
- evilelectron 6mo agoDaniel, your work is changing the world. More power to you. I setup a pipeline for inference with OCR, full text search, embedding and summarization of land records dating back 1800s. All powered by the GGUF's you generate and llama.cpp. People are so excited that they can now search the records in multiple languages that a 1 minute wait to process the document seems nothing. Thank you!
- danielhanchen 6mo agoOh appreciate it! Oh nice! That sounds fantastic! I hope Gemma-4 will make it even better! The small ones 2B and 4B are shockingly good haha!
- qingcharles 6mo agoJust switched from 3.1 Flash Lite to Gemma-4 31B on the AI Studio API since there is a generous 1500/day on non-billed projects. It's doing fantastic.
- polishdude20 6mo agoHey in really interested in your pipeline techniques. I've got some pdfs I need to get processed but processing them in the cloud with big providers requires redaction. Wondering if a local model or a self hosted one would work just as well.
- jorl17 6mo agoSeconded, would also love to hear your story if you would be willing
- evilelectron 6mo agoI run llama.cpp with Qwen3-VL-8B-Instruct-Q4_K_S.gguf with mmproj-F16.gguf for OCR and translation. I also run llama.cpp with Qwen3-Embedding-0.6B-GGUF for embeddings. Drupal 11 with ai_provider_ollama and custom provider ai_provider_llama (heavily derived from ai_provider_ollama) with PostreSQL and pgvector. People on site scan the documents and upload them for archival. The directory monitor looks for new files in the archive directories and once a new file is available, it is uploaded to Drupal. Once a new content is created in Drupal, Drupal triggers the translation and embedding process through llama.cpp. Qwen3-VL-8B is also used for chat and RAG. Client is familiar with Drupal and CMS in general and wanted to stay in a similar environment. If you are starting new I would recommend looking at docling.
- zaat 6mo agoThank you for your work. You have an answer on your page regarding "Should I pick 26B-A4B or 31B?", but can you please clarify if, assuming 24GB vRAM, I should pick a full precision smaller model or 4 bit larger model?
- danielhanchen 6mo agoThank you! I presume 24B is somewhat faster since it's only 4B activated - 31B is quite a large dense model so more accurate!
- ryandrake 6mo agoThis is one of the more confusing aspects of experimenting with local models as a noob. Given my GPU, which model should I use, which quantization of that model should I pick (unsloth tends to offer over a dozen!) and what context size should I use? Overestimate any of these, and the model just won't load and you have to trial-and-error your way to finding a good combination. The red/yellow/green indicators on huggingface.co are kind of nice, but you only know for sure when you try to load the model and allocate context.
- danielhanchen 6mo agoDefinitely Unsloth Studio can help - we recommend specific quants (like Gemma-4) and also auto calculate the context length etc!
- ryandrake 6mo agoWill have to try it out. I always thought that was more for fine-tuning and less for inference.
- danielhanchen 6mo agoOh yes sadly we partially mis-communicated haha - there's both and synthetic data generation + exporting!
- pentagrama 6mo agoHey, I tried to use Unsloth to run Gemma 4 locally but got stuck during the setup on Windows 11. At some point it asked me to create a password, and right after that it threw an error. Here’s a screenshot: https://imgur.com/a/sCMmqht https://imgur.com/a/sCMmqht This happened after running the PowerShell setup, where it installed several things like NVIDIA components, VS Code, and Python. At the end, PowerShell tell me to open a http://localhost URL in my browser, and that’s where I was prompted to set the password before it failed. Also, I noticed that an Unsloth icon was added to my desktop, but when I click it, nothing happens. For context, I’m not a developer and I had never used PowerShell before. Some of the steps were a bit intimidating and I wasn’t fully sure what I was approving when clicking through. The overall experience felt a bit rough for my level. It would be great if this could be packaged as a simple .exe or a standalone app instead of going through terminal and browser steps. Are there any plans to make something like that?
- danielhanchen 6mo agoApologies we just fixed it!! If you try again from source ie irm https://unsloth.ai/install.ps1 https://unsloth.ai/install.ps1 | iex it should work hopefully. If not - please at us on Discord and we'll help you! The Network error is a bummer - we'll check. And yes we're working on a .exe!!
- pentagrama 6mo agoIt worked! https://imgur.com/a/SOfiRhv https://imgur.com/a/SOfiRhv Thanks, will check it out tomorrow. Hope the unsloth-setup.exe > Windows App is coming soon! I think it will expand accessibility and user base.
- danielhanchen 6mo agoOh nice! Glad it worked! Yes!! We're working on the app!
- deleted 6mo ago[deleted]
- nnucera 6mo agoWow! Thank you very much!
- danielhanchen 6mo agoThanks!
- egeres 6mo agoThank you and your brother for all the amazing work, it's really inspiring to others <3
- danielhanchen 6mo agoThank you and appreciate it!
- jquery 6mo agoAwesome!! Thank you SO much for this.
- danielhanchen 6mo agoAppreciate it!
- Wowfunhappy 6mo agoHi! Do you ever make quants of the base models? I'm interested in experimenting with them in non-chat contexts.
- car 6mo agoYes, they are listed on huggingface. The instruction trained models have an 'it' in their name. https://huggingface.co/collections/unsloth/gemma-4 https://huggingface.co/collections/unsloth/gemma-4 Edit: Sorry, I'm not sure if this is a quant, but it says 'finetuned' from the Google Gemma 4 parent snapshot. It's the same size as the UD 8-bit quant though.
- Wowfunhappy 6mo agoOnly the 'it' models seem to have quants. I was really hoping to try a base model.
- kristjansson 6mo agoBasic quantization is easy if you have enough RAM (not VRAM) to load the weights.
- zobzu 6mo agoneat, time to update my spam filter model hehe
- danielhanchen 6mo agoHaha! Ye the model is really good
- Kye 6mo agoI haven't tried a local model in a while. I can only fit E4B in VRAM (8GB), but it's good enough that I can see it replacing Claude.ai for some things.
- akavel 6mo agoI'm trying to disable "thinking", but it doesn't seem to work (in llama.cpp). The usual `--reasoning-budget 0` doesn't seem to change it, nor `--chat-template-kwargs '{"enable_thinking":false}'` (both with `--jinja`). Am I missing something? EDIT: Ok, looks like there's yet another new flag for that in llama.cpp, and this one seems to work in this case: `--reasoning off`. FWIW, I'm doing some initial tries of unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL, and for writing some Nix, I'm VERY impressed - seems significantly better than qwen3.5-35b-a3b for me for now. Example commandline on a Macbook Air M4 32gb RAM: llama-cli -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL -t 1.0 --top-p 0.95 --top-k 64 -fa on --no-mmproj --reasoning-budget 0 -c 32768 --jinja --reasoning off (at release b8638, compiled with Nix)
- danielhanchen 6mo agoOh very cool! Will check the `--reasoning off` flag as well! Yep the models are really good!
- kapimalos 6mo agoNoob question. Why I would use this version over the original model?
- piyh 6mo ago1/3 the RAM & CPU consumed for 99% the performance
- trashcan2137 6mo agoand the EOS is "<turn|>". "<|channel>thought\n" is also used for the thinking trace! Can someone explain this to me? Why is this faux-XML important here?
- pertymcpert 6mo agoThat’s how the model is trained to signal the end to its generation and to indicate its thinking.
- sroussey 6mo agoThese are likely individual tokens. They are super common.
- genpfault 6mo agollama.cpp (b8642) auto-fits ~200k context on this 24GB RX 7900 XTX & it shows a solid 100+ tok/s ("S_TG t/s") on the first 32k of it, nice! ./llama-batched-bench -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \ -npp 1000,2000,4000,8000,16000,32000,64000,96000,128000 -ntg 128 -npl 1 -c 0 | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s | |-------|--------|------|--------|----------|----------|----------|----------|----------|----------| | 1000 | 128 | 1 | 1128 | 0.416 | 2404.87 | 1.064 | 120.29 | 1.480 | 762.20 | | 2000 | 128 | 1 | 2128 | 0.755 | 2649.86 | 1.075 | 119.04 | 1.830 | 1162.83 | | 4000 | 128 | 1 | 4128 | 1.501 | 2665.72 | 1.093 | 117.08 | 2.594 | 1591.49 | | 8000 | 128 | 1 | 8128 | 3.142 | 2545.85 | 1.114 | 114.87 | 4.257 | 1909.47 | | 16000 | 128 | 1 | 16128 | 6.908 | 2316.00 | 1.189 | 107.65 | 8.097 | 1991.73 | | 32000 | 128 | 1 | 32128 | 16.382 | 1953.31 | 1.278 | 100.12 | 17.661 | 1819.16 | | 64000 | 128 | 1 | 64128 | 43.427 | 1473.74 | 1.453 | 88.12 | 44.879 | 1428.89 | | 96000 | 128 | 1 | 96128 | 82.227 | 1167.50 | 1.623 | 78.86 | 83.850 | 1146.42 | |128000 | 128 | 1 | 128128 | 133.237 | 960.69 | 1.797 | 71.25 | 135.034 | 948.86 |
- danielhanchen 6mo agoOh nice that's pretty good!
- spwa4 6mo ago~50 tok/s on M1 Max 64Gb
- sillysaurusx 6mo agoTemperature 1.0 used to be bad for sampling. 0.7 was the better choice, and the difference in results were noticeable. You may want to experiment with this.
- danielhanchen 6mo agoYou might be right, but Google's recommendation was temp 1 etc primarily because all their benchmarks were used with these numbers, so it's better reproducibility for downstream tasks
- sillysaurusx 6mo agoFair, though putting a note in the readme about temperature 0.7 couldn't hurt. I wonder why they do benchmarks with 1 instead of 0.7... that's strange. 0.7 or 0.8 at most gives noticeably better samples.
- davedx 6mo agoReproducibility. They're benchmarks.
- sillysaurusx 6mo agoReproducibility is a matter of using the same input seeds, which jax can do. 0.7 vs 1.0 would make no difference for that. Without seeds, 0.7 would be less random than 1.0, so it'd be (slightly) more reproducible.
- sixhobbits 6mo agoThanks for this, I gave this guide to my Claude and he oneshot the unsloth and gemma4 set up on the old macbook he runs on. It's way faster than I expected, haven't tried out local models for a few generations but will be very nice when they become useful
- danielhanchen 6mo agoThanks! Oh nice! Ye local models are advancing much faster than I expected!
- zkmon 6mo agoHow does Gemma 4 26B A4B compare with Qwen3.5 35B A3B for same quants(4)
- deleted 6mo ago[deleted]
- rizzo94 6mo agoHuge fan of the Unsloth quants! Having reasoning and tool calling this accessible locally is a massive leap forward. The main hurdle I've found with local tool calling is managing the execution boundaries safely. I’ve started plugging these local models into PAIO to handle that. Since it acts as a hardened execution layer with strict BYOK sovereignty, it lets you actually utilize Gemma-4's tool calling capabilities without the low-level anxiety of a hallucination accidentally wiping your drive. It’s the perfect secure gateway for these advanced local models.
- mmaunder 6mo agoThis comment deserves it's own HN post. Thanks!
- jwr 6mo agoReally looking forward to testing and benchmarking this on my spam filtering benchmark. gemma-3-27b was a really strong model, surpassed later by gpt-oss:20b (which was also much faster). qwen models always had more variance.
- jeffbee 6mo agoDoes spam filtering really need a better model? My impression is that the whole game is based on having the best and freshest user-contributed labels.
- hrmtst93837 6mo ago[flagged]
- jeffbee 6mo agoIn my experience the contents of the message are all but totally irrelevant to the classification, and it is the behavior of the mailing peer that gives all the relevant features.
- mh- 6mo agoBased on how much blatant gmail->gmail spam I receive, the gmail team agrees with this strategy.
- drob518 6mo agoHe said it’s a benchmark.
- mhitza 6mo agoIf you wouldn't mind chatting about your usage, my email is in my profile, and I'd love to share experiences with other HNers using self-hosted models.
- a7om_com 6mo ago[flagged]
- flakiness 6mo agoIt's good they still have non-instruction-tuned models.
- minimaxir 6mo agoThe benchmark comparisons to Gemma 3 27B on Hugging Face are interesting: The Gemma 4 E4B variant (https://huggingface.co/google/gemma-4-E4B-it https://huggingface.co/google/gemma-4-E4B-it) beats the old 27B in every benchmark at a fraction of parameters. The E2B/E4B models also support voice input, which is rare.
- regularfry 6mo agoThinking vs non-thinking. There'll be a token cost there. But still fairly remarkable!
- DoctorOetker 6mo agoIs there a reason we can't use thinking completions to train non-thinking? i.e. gradient descent towards what thinking would have answered?
- joshred 6mo agoFrom what I've read, that's already part of their training. They are scored based on each step of their reasoning and not just their solution. I don't know if it's still the case, but for the early reasoning models, the "reasoning" output was more of a GUI feature to entertain the user than an actual explanation of the steps being followed.
- NitpickLawyer 6mo agoBest thing is that this is Apache 2.0 (edit: and they have base models available. Gemma3 was good for finetuning) The sizes are E2B and E4B (following gemma3n arch, with focus on mobile) and 26BA4 MoE and 31B dense. The mobile ones have audio in (so I can see some local privacy focused translation apps) and the 31B seems to be strong in agentic stuff. 26BA4 stands somewhere in between, similar VRAM footprint, but much faster inference.
- babelfish 6mo agoWow, 30B parameters as capable as a 1T parameter model?
- mhitza 6mo agoOn the above compared benchmarks is closer to other larger open weights models, and on par with GPT-OSS 120B, for which I also have a frame of reference.
- darshanmakwana 6mo agoThis is awesome! I will try to use them locally with opencode and see if they are usable inreplacement of claude code for basic tasks
- antirez 6mo agoFeaturing the ELO score as the main benchmark in chart is very misleading. The big dense Gemma 4 model does not seem to reach Qwen 3.5 27B dense model in most benchmarks. This is obviously what matters. The small 2B / 4B models are interesting and may potentially be better ASR models than specialized ones (not just for performances but since they are going to be easily served via llama.cpp / MLX and front-ends). Also interesting for "fast" OCR, given they are vision models as well. But other than that, the release is a bit disappointing.
- nabakin 6mo agoPublic benchmarks can be trivially faked. Lmarena is a bit harder to fake and is human-evaluated. I agree it's misleading for them to hyper-focus on one metric, but public benchmarks are far from the only thing that matters. I place more weight on Lmarena scores and private benchmarks.
- moffkalast 6mo agoLm arena is so easy to game that it's ceased to be a relevant metric over a year ago. People are not usable validators beyond "yeah that looks good to me", nobody checks if the facts are correct or not.
- nabakin 6mo agoIt's easy to game and human evaluation data has its trade-offs, but it's way easier to fake public benchmark results. I wish we had a source of high quality private benchmark results across a vast number of models like Lmarena. Having high quality human evaluation data would be a plus too.
- moffkalast 6mo agoWell there was this one [0] which is a black box but hasn't really been kept up to date with newer releases. Arguably we'd need lots of these since each one could be biased towards some use case or sell its test set to someone with more VC money than sense. [0] https://oobabooga.github.io/benchmark.html https://oobabooga.github.io/benchmark.html
- rvz 6mo agoOpen weight models once again marching on and slowly being a viable alternative to the larger ones. We are at least 1 year and at most 2 years until they surpass closed models for everyday tasks that can be done locally to save spending on tokens.
- echelon 6mo ago> We are at least 1 year and at most 2 years until they surpass closed models for everyday tasks that can be done locally to save spending on tokens. Until they pass what closed models today can do. By that time, closed models will be 4 years ahead. Google would not be giving this away if they believed local open models could win. Google is doing this to slow down Anthropic, OpenAI, and the Chinese, knowing that in the fullness of time they can be the leader. They'll stop being so generous once the dust settles.
- pixl97 6mo agoI mean, correct, but running open models locally will still massively drop your costs even if you still need to interface with large paid for models. Google will still make less money than if they were the only model that existed at the end of the day.
- ma2kx 6mo agoI think it will be less of a local versus cloud situation, but rather one where both complement each other. The next step will undoubtedly be for local LLMs to be fast and intelligent enough to allow for vocal conversation. A low-latency model will then run locally, enabling smoother conversations, while batch jobs in the cloud handle the more complex tasks. Google, at least, is likely interested in such a scenario, given their broad smartphone market. And if their local Gemma/Gemini-nano LLMs perform better with Gemini in the cloud, that would naturally be a significant advantage.
- jimbokun 6mo agoBut at that point, won’t there be very few tasks left where the average user can discern the difference in quality for most tasks?
- james2doyle 6mo agoHmm just tried the google/gemma-4-31B-it through HuggingFace (inference provider seems to be Novita) and function/tool calling was not enabled...
- james2doyle 6mo agoYeah you can see here that tool calling is disabled: https://huggingface.co/inference/models?model=google%2Fgemma-4-31B-it https://huggingface.co/inference/models?model=google%2Fgemma... At least, as of this post
- linolevan 6mo agoHosted on Parasail + Google (both for free, as of now) themselves, probably would give those a shot
- hyjohnnychin 6mo agoTool calling is enabled now
- originalvichy 6mo agoThe wait is finally over. One or two iterations, and I’ll be happy to say that language models are more than fulfilling my most common needs when self-hosting. Thanks to the Gemma team!
- adamtaylor_13 6mo agoWhat sort of tasks are you using self-hosting for? Just curious as I've been watching the scene but not experimenting with self-hosting.
- irishcoffee 6mo agoI would personally be much more interested in using LLMs if I didn’t need to depend on an internet connection and spending money on tokens.
- vunderba 6mo agoNot OP but one example is that recent VL models are more than sufficient for analyzing your local photo albums/images for creating metadata / descriptions / captions to help better organize your library.
- kejaed 6mo agoAny pointers on some local VLMs to start with?
- canyon289 6mo agoYou could try Gemma4 :D
- vunderba 6mo agoThe easiest way to get started is probably to use something like Ollama and use the `qwen3-vl:8b` 4‑bit quantized model [1]. It's a good balance between accuracy and memory, though in my experience, it's slower than older model architectures such as Llava. Just be aware Qwen-VL tends to be a bit verbose [2], and you can’t really control that reliably with token limits - it'll just cut off abruptly. You can ask it to be more concise but it can be hit or miss. What I often end up doing and I admit it's a bit ridiculous is letting Qwen-VL generate its full detailed output, and then passing that to a different LLM to summarize. - [1] https://ollama.com/library/qwen3-vl:8b https://ollama.com/library/qwen3-vl:8b - [2] https://mordenstar.com/other/vlm-xkcd https://mordenstar.com/other/vlm-xkcd
- scrlk 6mo agoComparison of Gemma 4 vs. Qwen 3.5 benchmarks, consolidated from their respective Hugging Face model cards: | Model | MMLUP | GPQA | LCB | ELO | TAU2 | MMMLU | HLE-n | HLE-t | |----------------|-------|-------|-------|------|-------|-------|-------|-------| | G4 31B | 85.2% | 84.3% | 80.0% | 2150 | 76.9% | 88.4% | 19.5% | 26.5% | | G4 26B A4B | 82.6% | 82.3% | 77.1% | 1718 | 68.2% | 86.3% | 8.7% | 17.2% | | G4 E4B | 69.4% | 58.6% | 52.0% | 940 | 42.2% | 76.6% | - | - | | G4 E2B | 60.0% | 43.4% | 44.0% | 633 | 24.5% | 67.4% | - | - | | G3 27B no-T | 67.6% | 42.4% | 29.1% | 110 | 16.2% | 70.7% | - | - | | GPT-5-mini | 83.7% | 82.8% | 80.5% | 2160 | 69.8% | 86.2% | 19.4% | 35.8% | | GPT-OSS-120B | 80.8% | 80.1% | 82.7% | 2157 | -- | 78.2% | 14.9% | 19.0% | | Q3-235B-A22B | 84.4% | 81.1% | 75.1% | 2146 | 58.5% | 83.4% | 18.2% | -- | | Q3.5-122B-A10B | 86.7% | 86.6% | 78.9% | 2100 | 79.5% | 86.7% | 25.3% | 47.5% | | Q3.5-27B | 86.1% | 85.5% | 80.7% | 1899 | 79.0% | 85.9% | 24.3% | 48.5% | | Q3.5-35B-A3B | 85.3% | 84.2% | 74.6% | 2028 | 81.2% | 85.2% | 22.4% | 47.4% | MMLUP: MMLU-Pro GPQA: GPQA Diamond LCB: LiveCodeBench v6 ELO: Codeforces ELO TAU2: TAU2-Bench MMMLU: MMMLU HLE-n: Humanity's Last Exam (no tools / CoT) HLE-t: Humanity's Last Exam (with search / tool) no-T: no think
- kpw94 6mo agoWild differences in ELO compared to tfa's graph: https://storage.googleapis.com/gdm-deepmind-com-prod-public/media/original_images/QInH6awnEGY0Anki/gemma__gemma-4__elo-size__dark_QICdcNx.svg https://storage.googleapis.com/gdm-deepmind-com-prod-public/... (Comparing Q3.5-27B to G4 26B A4B and G4 31B specifically) I'd assume Q3.5-35B-A3B would performe worse than the Q3.5 deep 27B model, but the cards you pasted above, somehow show that for ELO and TAU2 it's the other way around... Very impressed by unsloth's team releasing the GGUF so quickly, if that's like the qwen 3.5, I'll wait a few more days in case they make a major update. Overall great news if it's at parity or slightly better than Qwen 3.5 open weights, hope to see both of these evolve in the sub-32GB-RAM space. Disappointed in Mistral/Ministral being so far behind these US & Chinese models
- coder543 6mo ago
- ceroxylon 6mo agoEven with search grounding, it scored a 2.5/5 on a basic botanical benchmark. It would take much longer for the average human to do a similar write-up, but they would likely do better than 50% hallucination if they had access to a search engine.
- WarmWash 6mo agoEven multimodal models are still really bad when it comes to vision. The strength is still definitely language.
- nostrebored 6mo agoTraining for tasks still works petty well, but “vision” is a super broad domain and most seem optimized for OCR and screen processing (which have verifiable outputs and relatively straightforward data generation)
- wg0 6mo agoGoogle might not have the best coding models (yet) but they seem to have the most intelligent and knowledgeable models of all especially Gemini 3.1 Pro is something. One more thing about Google is that they have everything that others do not: 1. Huge data, audio, video, geospatial 2. Tons of expertise. Attention all you need was born there. 3. Libraries that they wrote. 4. Their own data centers and cloud. 4. Most of all, their own hardware TPUs that no one has. Therefore once the bubble bursts, the only player standing tall and above all would be Google.
- chasd00 6mo agoNot sure why you're being downvoted, the other thing Google has is Google. They just have to spend the effort/resources to keep up and wait for everyone else to go bankrupt. At the end of the day I think Google will be the eventual LLM winner. I think this is why Meta isn't really in the race and just releases open weight models, the writing is on the wall. Also, probably why Apple went ahead and signed a deal with Google and not OpenAI or Anthropic.
- wg0 6mo agoI don't know why I am downvoted but Google has data, expertise, hardware and deep pockets. This whole LLM thing is invented at Google and machine learning ecosystem libraries come from Google. I don't know how people can be so irrational discounting Google's muscle. Others have just borrowed data, money, hardware and they would run out of resources for sure.
- greenavocado 6mo agoThis remains true so long as advertisers give Google money.
- bitpush 6mo agoWhy wouldnt advertisers give Google money? Are you noticing any shift in trend?
- 6mo ago
- mudkipdev 6mo agoCan't wait for gemma4-31b-it-claude-opus-4-6-distilled-q4-k-m on huggingface tomorrow
- entropicdrifter 6mo agoI'd rather see a distill on the 26B model that uses only 3.8B parameters at inference time. Seems like it will be wildly productive to use for locally-hosted stuff
- indrora 6mo agogemma4-31b-it-claude-opus-4-6-distilled-abliterated-heretic-GGUF-q4-k-m
- bertili 6mo agoQwen: Hold my beer https://news.ycombinator.com/item?id=47615002 https://news.ycombinator.com/item?id=47615002
- xfalcox 6mo agoComparing a model you can downloads weights for with an API-only model doesn't make much sense.
- regularfry 6mo agoMy money's on whatever models qwen does release edging ahead. Probably not by much, but I reckon they'll be better coders just because that's where qwen's edge over gemma has always been. Plus after having seen this land they'll probably tack on a couple of epochs just to be sure.
- svachalek 6mo agoThe Qwen Plus models should be compared to Gemini, not Gemma.
- fooker 6mo agoWhat's a realistic way to run this locally or a single expensive remote dev machine (in a vm, not through API calls)?
- matja 6mo agoI'm running Gemma 4 with the llama.cpp web UI. https://unsloth.ai/docs/models/gemma-4 https://unsloth.ai/docs/models/gemma-4 > Gemma 4 GGUFs > "Use this model" > llama.cpp > llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q8_0 If you already have llama.cpp you might need to update it to support Gemma 4.
- deleted 6mo ago[deleted]
- evanbabaallos 6mo ago[flagged]
- deleted 6mo ago[deleted]
- heraldgeezer 6mo agoGemma vs Gemini? I am only a casual AI chatbot user, I use what gives me the most and best free limits and versions.
- daemonologist 6mo agoGemma will give you the most, Gemini will give you the best. The former is much smaller and therefore cheaper to run, but less capable. Although I'm not sure whether Gemma will be available even in aistudio - they took the last one down after people got it to say/do questionable stuff. It's very much intended for self-hosting.
- BoorishBears 6mo agoWell specifically a congressperson got it to hallucinate stuff about them then wrote an agry letter But I checked and it's there... but in the UI web search can't be disabled (presumably to avoid another egg on face situation)
- worldsavior 6mo agoGemma is only 10s of billion parameters, Gemini is 100s.
- deleted 6mo ago[deleted]
- VadimPR 6mo agoGemma 3 E4E runs very quick on my Samsung S26, so I am looking forward to trying Gemma 4! It is fantastic to have local alternatives to frontier models in an offline manner.
- snthpy 6mo agoWhat's the easiest way to install these on an Android phone/Samsung?
- nolist_policy 6mo agoGoogle AI Edge Gallery: https://github.com/google-ai-edge/gallery/releases https://github.com/google-ai-edge/gallery/releases
- VadimPR 6mo agoI use LM Studio, but there's a comment here offering another tool as well.
- VadimPR 6mo agoLM Playground, not LM Studio.
- canyon289 6mo agoHi all! I work on the Gemma team, one of many as this one was a bigger effort given it was a mainline release. Happy to answer whatever questions I can
- wahnfrieden 6mo agoHow is the performance for Japanese, voice in particular?
- canyon289 6mo agoI dont have the metrics off hand, but I'd say try it and see if you're impressed! What matters at the end of the day is if its useful for your use cases and only you'll be able to assess that!
- k3nz0 6mo agoHow do you test codeforces ELO?
- canyon289 6mo agoOn this one I dont know :) I'll ask my friends on the evaluation side of things how they do this
- azinman2 6mo agoHow do the smaller models differ from what you guys will ultimately ship on Pixel phones? What's the business case for releasing Gemma and not just focusing on Gemini + cloud only?
- canyon289 6mo agoIts hard to say because Pixel comes prepacked with a lot of models, not just ones that that are text output models. With the caveat that I'm not on the pixel team and I'm not building _all_ the models that are on google's devices, its evident there are many models that support the Android experience. For example the one mentioned here https://store.google.com/us/magazine/magic-editor?hl=en-US&pli=1 https://store.google.com/us/magazine/magic-editor?hl=en-US&p...
- mwizamwiinga 6mo ago[flagged]
- chrislattner 6mo agoIf you want the fastest open source implementation on Blackwell and AMD MI355, check out Modular's MAX nightly. You can pip install it super fast, check it out here: https://www.modular.com/blog/day-zero-launch-fastest-performance-for-gemma-4-on-nvidia-and-amd?utm_campaign=day0&utm_source=hn_chris https://www.modular.com/blog/day-zero-launch-fastest-perform... -Chris Lattner (yes, affiliated with Modular :-)
- nabakin 6mo agoFaster than TensorRT-LLM on Blackwell? Or do you not consider TensorRT-LLM open source because some dependencies are closed source?
- melodyogonna 6mo agoI reviewed the TensorRT-LLM commit history from the past few days and couldn't find any updates regarding Gemma 4 support. By contrast, here is the reference for MAX:https://github.com/modular/modular/commit/57728b23befed8f3b44e24d67d41b275bd3e49d4 https://github.com/modular/modular/commit/57728b23befed8f3b4...
- nabakin 6mo agoIf OP meant they have the fastest implementation of Gemma 4 on Blackwell at the moment, I guess that is technically true. I doubt that will hold up when TensorRT-LLM finishes their implementation though.
- simonw 6mo agoI ran these in LM Studio and got unrecognizable pelicans out of the 2B and 4B models and an outstanding pelican out of the 26b-a4b model - I think the best I've seen from a model that runs on my laptop. https://simonwillison.net/2026/Apr/2/gemma-4/ https://simonwillison.net/2026/Apr/2/gemma-4/ The gemma-4-31b model is completely broken for me - it just spits out "---\n" no matter what prompt I feed it. I got a pelican out of it via the AI Studio API hosted model instead.
- wordpad 6mo agoDo you think it's just part of their training set now?
- simonw 6mo agoIf it's part of their training set why do the 2B and 4B models produce such terrible SVGs?
- vessenes 6mo agoWe were promised full SVG zoos, Simon. I want to see SVG pangolins please
- retinaros 6mo agobecause generating nice looking svg requires handling code, shapes, long context, reasoning and at 2b you most likely will break the syntax of the file 9 times out of 10 if you train for that. or you will need to go for simpler pelicans. might not be worth to ft on a 2b. but on their top tier open model it is definitly worth it. even not directly but just crawling a github would make it train on your pelicans.
- wolttam 6mo agoBecause it is in their training set but it's unrealistic to expect a 2B or 4B model to be able to perfectly reproduce everything it's seen before. The training no doubt contributed to their ability to (very) loosely approximate an SVG of pelican on a bicycle, though. Frankly I'm impressed
- aplomb1026 6mo ago[dead]
- bertili 6mo agoThe timing is interesting as Apple supposedly will distill google models in the upcoming Siri update [1]. So maybe Gemma is a lower bound on what we can expect baked into iPhones. [1] https://news.ycombinator.com/item?id=47520438 https://news.ycombinator.com/item?id=47520438
- whhone 6mo agoThe LiteRT-LM CLI (https://ai.google.dev/edge/litert-lm/cli https://ai.google.dev/edge/litert-lm/cli) provides a way to try the Gemma 4 model. # with uvx uvx litert-lm run \ --from-huggingface-repo=litert-community/gemma-4-E2B-it-litert-lm \ gemma-4-E2B-it.litertlm
- DeepYogurt 6mo agomaybe a dumb question but what what does the "it" stand for in the 31B-it vs 31B?
- bigyabai 6mo agoInstruction Tuned. It indicates that thinking tokens (eg <think> </think>) are not included in training.
- deleted 6mo ago[deleted]
- flux3125 6mo agoThat’s not what it means. "-it" just indicates the model is instruction-tuned, i.e. trained to follow prompts and behave like an assistant. It doesn’t imply anything about whether thinking tokens like <think>....</think> were included or excluded during training. Thats a separate design choice and varies by model.
- DeepYogurt 6mo agoWhat does that mean for a user of the model? Is the "-it" version more direct with solutions or something?
- nolist_policy 6mo agoUse the it versions. The other versions are base models without post-training. E.g. base models are trained to regurgitate raw wikipedia, books, etc. Then these base models are post-trained into instruction-tuned models where they learn to act as a chat assistant.
- petu 6mo agoIt means that model was tuned to to act as chat bot. So write a reply on behalf of assistant and stop generating (by inserting special "end of turn" token to signal inference engine to stop generation). Base model (without instruction/chat tuning) just generates text non stop ("autocomplete on steroids") and text is not necessarily even formatted as chat -- most text in training data isn't dialogue, after all.
- virgildotcodes 6mo agoDownloaded through LM Studio on an M1 Max 32GB, 26B A4B Q4_K_M First message: https://i.postimg.cc/yNZzmGMM/Screenshot-2026-04-03-at-12-44-38-AM.png https://i.postimg.cc/yNZzmGMM/Screenshot-2026-04-03-at-12-44... Not sure if I'm doing something wrong? This more or less reflects my experience with most local models over the last couple years (although admittedly most aren't anywhere near this bad). People keep saying they're useful and yet I can't get them to be consistently useful at all.
- solarkraft 6mo agoWow, just like its larger brother! I had a similarly bad experience running Qwen 3.5 35b a3b directly through llama.cpp. It would massively overthink every request. Somehow in OpenCode it just worked. I think it comes down to temperature and such (see daniel‘s post), but I haven’t messed with it enough to be sure.
- flux3125 6mo agoYou're not doing anything wrong, that's expected
- sigbottle 6mo agoThere are so many heavy hitting cracked people like daniel from unsloth and chris lattner coming out of the woodworks for this with their own custom stuff. How does the ecosystem work? Have things converged and standardized enough where it's "easy" (lol, with tooling) to swap out parts such as weights to fit your needs? Do you need to autogen new custom kernels to fix said things? Super cool stuff.
- bredren 6mo agoThanks for the notes, for those interested in learning more: - Lattner tweeted a link to this: https://www.modular.com/blog/day-zero-launch-fastest-performance-for-gemma-4-on-nvidia-and-amd https://www.modular.com/blog/day-zero-launch-fastest-perform... - Unsloth prior post on gemma 3 finetuning: https://unsloth.ai/blog/gemma3 https://unsloth.ai/blog/gemma3
- einpoklum 6mo agoD: Di Gi Charat does not like this nyo! Gemma is supposed to help Dejiko-chan nyo! G: They offered a very compelling benefits package gemma!
- karimf 6mo agoI'm curious about the multimodal capabilities on the E2B and E4B and how fast is it. In ChatGPT right now, you can have a audio and video feed for the AI, and then the AI can respond in real-time. Now I wonder if the E2B or the E4B is capable enough for this and fast enough to be run on an iPhone. Basically replicating that experience, but all the computations (STT, LLM, and TTS) are done locally on the phone. I just made this [0] last week so I know you can run a real-time voice conversation with an AI on an iPhone, but it'd be a totally different experience if it can also process a live camera feed. https://github.com/fikrikarim/volocal https://github.com/fikrikarim/volocal
- functional_dev 6mo agoyeah, it appears to support audio and image input.. and runs on mobile devices with 256K context window!
- coder543 6mo agoThe E2B and E4B models support 128k context, not 256k, and even with the 128k... it could take a long time to process that much context on most phones, even with the processor running full tilt. It's hard to say without benchmarks, but 128k supported isn't the same as 128k practical. It will be interesting to see.
- fy20 6mo agoI just want to say thanks. Finding out about these kind of projects that people are working on is what I come to HN for, and what excites me about software engineering!
- karimf 6mo agoThank you for the kind words!
- karimf 6mo agoUpdate: Just made one that runs on Macbook M3 Pro https://github.com/fikrikarim/parlor https://github.com/fikrikarim/parlor
- Analog24 6mo agoSo the "E2B" and "E4B" models are actually 5B and 8B parameters. Are we really going to start referring to the "effective" parameter count of dense models by not including the embeddings? These models are impressive but this is incredibly misleading. You need to load the embeddings in memory along with the rest of the model so it makes no sense o exclude them from the parameter count. This is why it actually takes 5GB of RAM to run the "2B" model with 4-bit quantization according to Unsloth (when I first saw that I knew something was up).
- nolist_policy 6mo agoThese are based on the Gemma 3n architecture so E2B only needs 2Gb for text2text generation: https://ai.google.dev/gemma/docs/gemma-3n#parameters https://ai.google.dev/gemma/docs/gemma-3n#parameters You can think of the per layer-embeddings as a vector database so you can in theory serve it directly from disk.
- bearjaws 6mo agoThe labels on the table read "Gemma 431B IT" which reads as 431B parameter model, not Gemma 4 - 31B...
- kuboble 6mo agoIm really looking forward to trying it out. Gemma 3 was the first model that I have liked enough to use a lot just for daily questions on my 32G gpu.
- stephbook 6mo agoKind of sad they didn't release stronger versions. $dayjob offers strong NVidias that are hungry for models and are stuck running llama, gpt-oss etc. Seems like Google and Anthropic (which I consider leaders) would rather keep their secret sauce to themselves – understandable.
- matt765 6mo agoI'll wait for the next iteration
- stevenhubertron 6mo agoStill pretty unusable on Raspberry Pi 5, 16gb despite saying its built for it, from the E4B model total duration: 12m41.34930419s load duration: 549.504864ms prompt eval count: 25 token(s) prompt eval duration: 309.002014ms prompt eval rate: 80.91 tokens/s eval count: 2174 token(s) eval duration: 12m36.577002621s eval rate: 2.87 tokens/s Prompt: whats a great chicken breast recipe for dinner tonight?
- stevenhubertron 6mo agoOn my MBP M4 Pro 48gb same model/question while multitasking with Figma, email etc: total duration: 37.44872875s load duration: 145.783625ms prompt eval count: 25 token(s) prompt eval duration: 215.114666ms prompt eval rate: 116.22 tokens/s eval count: 1989 token(s) eval duration: 36.614398076s eval rate: 54.32 tokens/s
- daveguy 6mo agoFyi, it took me a while to find the meaning of the "-it" in some models. That's how Google designates "instruction tuned". Come on Google. Definite your acronyms.
- 0xbadcafebee 6mo agoGemma 3 models were pretty bad, so hopefully they got Gemma 4 to at least come close to the other major open weights
- nolist_policy 6mo agoBad at coding. Good for everything else.
- DanDeBugger 6mo ago[dead]
- hikarudo 6mo agoAlso checkout Deepmind's "The Gemma 4 Good Hackathon" on kaggle: https://www.kaggle.com/competitions/gemma-4-good-hackathon https://www.kaggle.com/competitions/gemma-4-good-hackathon
- swalsh 6mo agoI gave the same prompt (a small rust project that's not easy, but not overly sophisticated) to both Gemma-4 26b and Qwen 3.5 27b via OpenCode. Qwen 3.5 ran for a bit over an hour before I killed it, Gemma 4 ran for about 20 minutes before it gave up. Lots of failed tool calls. I asked codex to write a summary about both code bases. "Dev 1" Qwen 3.5 "Dev 2" Gemma 4 Dev 1 is the stronger engineer overall. They showed better architectural judgment, stronger completeness, and better maintainability instincts. The weakness is execution rigor: they built more, but didn’t verify enough, so important parts don’t actually hold up cleanly. Dev 2 looks more like an early-stage prototyper. The strength is speed to a rough first pass, but the implementation is much less complete, less polished, and less dependable. The main weakness is lack of finish and technical rigor. If I were choosing between them as developers, I’d take Dev 1 without much hesitation. Looking at the code myself, i'd agree with codex.
- coder543 6mo agoThere are issues with the chat template right now[0], so tool calling does not work reliably[1]. Every time people try to rush to judge open models on launch day... it never goes well. There are ~always bugs on launch day. [0]: https://github.com/ggml-org/llama.cpp/pull/21326 https://github.com/ggml-org/llama.cpp/pull/21326 [1]: https://github.com/ggml-org/llama.cpp/issues/21316 https://github.com/ggml-org/llama.cpp/issues/21316
- emidoots 6mo agowas just merged
- coder543 6mo agoIt was just an example of a bug, not that it was the only bug. I’ve personally reported at least one other for Gemma 4 on llama.cpp already. In a few days, I imagine that Gemma 4 support should be in better shape.
- stavros 6mo agoWhat causes these? Given how simple the LLM interface is (just completion), why don't teams make a simple, standardized template available with their model release so the inference engine can just read it and work properly? Can someone explain the difficulty with that?
- gunalx 6mo agoWe didnt get deepseek v4, but gemma 4. Cant complain.
- Deegy 6mo agoSo what's the business strategy here? Google is the only USA based frontier lab releasing open models. I know they aren't doing it out of the goodness of their hearts.
- artificialprint 6mo agoRelease open weights so competitors can't raise good money, then rear naked choke when they run dry
- robocat 6mo agoUsing Brazilian Jiu-Jitsu (BJJ) technical terms is confusing. Sports allusions don't travel well between cultures, especially if they sound seedy.
- golfer 6mo agoI found it plucky and intriguing. A great metaphor, not often seen in tech. Not everything has to be in the lowest common denominator of language.
- g947o 6mo agohttps://openai.com/index/introducing-gpt-oss/ https://openai.com/index/introducing-gpt-oss/
- mchusma 6mo agoFor those curious, on openrouter this is $0.14 input and $0.40 output, or ballpark half of Gemini flash lite 3.1 (googles current cheapest current gen closed model)
- mchusma 6mo agoDoing a bit more research, this looks like it might perform roughly as well on text tasks with modest context windows, so may be just a better cheaper option unless you need a million token window.
- bibimsz 6mo agois it good? what's it good for?
- neonstatic 6mo agoPrompt: > what is the Unix timestamp for this: 2026-04-01T16:00:00Z Qwen 3.5-27b-dwq > Thought for 8 minutes 34 seconds. 7074 tokens. > The Unix timestamp for 2026-04-01T16:00:00Z is: > 1775059200 (my comment: Wednesday, 1 April 2026 at 16:00:00) Gemma-4-26b-a4b > Thought for 33.81 seconds. 694 tokens. > The Unix timestamp for 2026-04-01T16:00:00Z is: > 1775060800 (my comment: Wednesday, 1 April 2026 at 16:26:40) Gemma considered three options to solve this problem. From the thinking trace: > Option A: Manual calculation (too error-prone). > Option B: Use a programming language (Python/JavaScript). > Option C: Knowledge of specific dates. It then wrote a python script: from datetime import datetime, timezone date_str = "2026-04-01T16:00:00Z" # Replace Z with +00:00 for ISO format parsing or just strip it dt = datetime.strptime(date_str, "%Y-%m-%dT%H:%M:%SZ").replace(tzinfo=timezone.utc) ts = int(dt.timestamp()) print(ts) Then it verified the timestamp with a command: date -u -d @1775060800 All of this to produce a wrong result. Running the python script it produced gives the correct result. Running the verification date command leads to a runtime error (hallucinated syntax). On the other hand Qwen went straight to Option A and kept overthinking the question, verifying every step 10 times, experienced a mental breakdown, then finally returned the right answer. I think Gemma would be clearly superior here if it used the tools it came up with rather than hallucinating using them.
- augusto-moura 6mo agoThe date command is not wrong, it works on GNU date, if you are in MacOS try running gdate instead (if it is installed): gdate -u -d @1775060800 To install gdate and GNU coreutils: brew install coreutils The date command still prints the incorrect value: Wed Apr 1 16:26:40 UTC 2026
- neonstatic 6mo agoGood catch, I just ran it verbatim in iTerm2 on macOs: date -u -d @1775060800 date: illegal option -- d btw. how do you format commands in a HN comment correctly?
- augusto-moura 6mo ago
- popinman322 6mo agoDoes anyone know whether we'll be receiving transcoders for this batch of models? We got them for Gemma 3, but maybe that was a one-off.
- wei03288 6mo ago[dead]
- kelsey98765431 6mo ago[dead]
- d4rkp4ttern 6mo agoFor token-generation speed, a challenging test is to see how it performs in a code-agent harness like Claude Code, which has anywhere between 15-40K tokens from the system prompt itself (+ tools/skills etc). Here the 26B-A4B variant is head and shoulders above recent open-weight models, at least on my trusty M1 Max 64GB MacBook. I set up Claude Code to use this variant via llama-server, with 37K tokens initial context, and it performs very well: ~40 tokens/sec, far better than Qwen3.5-35B-A3B, though I don't know yet about the intelligence or tool-calling consistency. Prompt processing speed is comparable to the Qwen variant at ~400 tok/s. My informal tests, all with roughly 30K-37K tokens initial context: ┌────────────────────┬───────────────┬────────────┐ │ Model │ Active Params │ tg (tok/s) │ ├────────────────────┼───────────────┼────────────┤ │ Gemma-4-26B-A4B │ 4B │ ~40 │ ├────────────────────┼───────────────┼────────────┤ │ GPT-OSS-20B │ 3.6B │ ~17-38 │ ├────────────────────┼───────────────┼────────────┤ │ Qwen3-30B-A3B │ 3B │ ~15-27 │ ├────────────────────┼───────────────┼────────────┤ │ GLM-4.7-Flash │ 3B │ ~12-13 │ ├────────────────────┼───────────────┼────────────┤ │ Qwen3.5-35B-A3B │ 3B │ ~12 │ ├────────────────────┼───────────────┼────────────┤ │ Qwen3-Next-80B-A3B │ 3B │ ~3-5 │ └────────────────────┴───────────────┴────────────┘ Full instructions for running this and other open-weight models with Claude Code are here: https://pchalasani.github.io/claude-code-tools/integrations/local-llms/#gemma-4-26b-a4b--google-moe-with-vision https://pchalasani.github.io/claude-code-tools/integrations/...
- JoshPurtell 6mo agogpt oss 20b is not dense
- d4rkp4ttern 6mo agoThanks, fixed
- janalsncm 6mo agoI don’t think this should be dead @dang?
- fc417fc802 6mo agoIt's no longer dead (I vouched) or you couldn't have replied. Also handles don't work here you have to email.
- stefs 6mo agoi get a lot of tool call errors with gemma-4-26b-a4b, because the tokens don't seem to match up.
- synergy20 6mo agoa dumb question, is this better than qwen3.5 and I thus should switch over?
- AnonyMD 6mo agoIt's great that it can run in a local environment.
- yalogin 6mo agoDo these come in quantized variants too? I mean may be 10B or lower? Wonder how they function.
- screenshotapi 6mo agoI love how they have both the 31B dense and 26B MoE, both fit well locally. Any MLX ports already?
- try-working 6mo agoThe biggest story here is that this is Google handing Qwen the SOTA crown for small and medium models. For the first time ever, a Chinese lab is at the frontier. Google and Nvidia are significantly behind, not just on benchmarks but real-world performance like tool calling accuracy.
- vicchenai 6mo agoThe 4B being this capable is honestly surprising. Ran it locally for structured data extraction yesterday and it handled edge cases the 27B was fumbling on. Didn't expect to swap down that fast.
- vigneshj 6mo agoGreat one to have
- simonw 6mo agoAnyone figured out a recipe to run Gemma 4 E2B or E4B against audio files locally on a Mac?
- coder543 6mo agoIf you search the model card[0], there is a section titled "Code for processing Audio", which you can probably use to test things out. But, the model card makes the audio support seem disappointing: > Audio supports a maximum length of 30 seconds. [0]: https://huggingface.co/google/gemma-4-26B-A4B-it#getting-started https://huggingface.co/google/gemma-4-26B-A4B-it#getting-sta...
- rahimnathwani 6mo agoPrince Canuma just updated mlx-vlm: https://x.com/i/status/2039815307821199709 https://x.com/i/status/2039815307821199709 So something like this should work: https://x.com/i/status/1938328542699503723 https://x.com/i/status/1938328542699503723
- devnotes77 6mo ago[dead]
- noritaka88 6mo ago[flagged]
- gslepak 6mo ago"casually dropping the most capable open weights on the planet" — @RyanMullins Google folks do something really cool! Gemma4 source: https://github.com/huggingface/transformers/pull/45192 https://github.com/huggingface/transformers/pull/45192
- kvntrnz 6mo agoLet's gooo keen to try it out
- Retro_Dev 6mo agoI'm very pleased with the performance of the largest gemma4 model (which I tested through ollama). My singular data point on whether an LLM remembers things well is whether it can translate toki pona to (and from) English. I find it easy to evaluate because I know the language. This local LLM marks the first version that 1) doesn't hallucinate words - at least, for the largest model - and 2) uses common word-phrases that other toki pona speakers use, and most importantly 3) can actually run on my laptop.
- curioussquirrel 6mo agoWe're doing multilingual testing and I can confirm what you've observed: Gemma 4 is surprisingly good at multilingual tasks, especially given its size. This is mostly true for the dense 31B model.
- Reubend 6mo agoI would suggest that people stop overfocusing on benchmarks, and give this a try. Gemma 4 is performing really well for me, and seems to hallucinate much less than other models I tried in this size range.
- burgerquizz 6mo agoI want to embed a lightweight local model to be used for my webapp to use it without thinking about token price. is there an acceptable way to do it today?
- mybigbro 6mo agoWent through the official blog and the developers post, no mention of TurboQuant anywhere. Google's own research team tested it on Gemma models for KV-cache compression to 3 bits, so it's surprising it's not mentioned in this release. Anyone know if it's baked in already or if we'd need to apply it ourselves? Would love to run the 26B MoE locally as a daily driver.
- i386 6mo agoYou can try this new model live using mesh-llm right now: https://www.anarchai.org/dashboard https://www.anarchai.org/dashboard
- chrischavez 6mo agoWent through the official blog and the developers post, no mention of TurboQuant anywhere. Google's own research team tested it on Gemma models for KV-cache compression to 3 bits, so it's surprising it's not mentioned in this release. Anyone know if it's baked in already or if we'd need to apply it ourselves? Would love to run the 26B MoE locally as a daily driver.
- nl 6mo agoGemma-4-E4B-it scored 15/25 on my https://sql-benchmark.nicklothian.com/#all-data https://sql-benchmark.nicklothian.com/#all-data (agentic SQL generation). The naming is a bit odd - E4B is "4.5B effective, 8B with embeddings", so despite the name it is probably best compared with the 8B/9B class models and is competitive with them. Qwen3.5-9B also scores 15/25 in thinking mode for example. The best 9B model I've found is Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2 which gets to 17/25 gemma-4-E2B (4bit quant) scored 12/25, but is really a 5B model. That's the same as NVIDIA-Nemotron-3-Nano-4B which is the best 4B model I've found (yes, better than Qwen 4B). That's a great score for a small model.
- alecthomas 6mo agoOh this page is great! I just released AIM [1] which is a tool that generates verified SQL migrations using LLMs, and I tested a bunch of models manually. I think I'll just link to your page too! [1] https://github.com/alecthomas/aim https://github.com/alecthomas/aim
- GaggiX 6mo ago>so despite the name it is probably best compared with the 8B/9B It runs much faster than a standard 8B/9B model, the name is given by the fact that it uses per-layer embedding (PLE).
- neonstatic 6mo agoVery happy to see updates to your benchmark. Looking forward to inclusion of larger Gemma 4 models!
- nl 6mo agoThe medium one on OpenRouter didn't support tools when I tried it. I will update when there is one.
- chromatin 6mo agoI love that you are doing this test. However, as it purports to be a test of "English-to-SQL", your hardest question (Q9) seems ungrammatical: > Show order lines, revenue, units sold, revenue per unit (total revenue ÷ total units sold), average list price per product in the subcategory, gross profit, and margin percentage for each product subcategory. In particular, the clause "in the subcategory, gross profit, and margin percentage for each product subcategory" is ambiguous, and I wonder if more models would pass if the English were reformulated to be correct. (it's also notable that Claude Opus 4.6 and Sonnet 4.6 both "missed" this one)
- aggregator-ios 6mo agoI tested the E2B and E4B models and they get close but inaccurate (non working) results when generating jq queries from natural language. This is of importance to me as I work on https://jsonquery.app https://jsonquery.app and would prefer to use a model that works well with browser inference. gemma-4-26b-a4b-it and gemma-4-31b-it produced accurate results in a few of my tests. But those are 50-60GB in size. Chrome has a developer preview that bundles Gemini Nano (under 2GB) and it used to work really well, but requires a few switches to be manually switched on, and has recently gotten worse in quality when testing for jq generation.
- curioussquirrel 6mo agoSame, I quickly tested it for code gen and it produced mostly good code for simple problems, but it sometimes hallucinated words in non-English scripts inside the code.
- Agent01001 6mo agolooks cool
- davedyarrow 6mo ago[dead]
- ahwgiuwh 6mo ago[flagged]
- ahwgiuwh 6mo ago[flagged]
- ahwg1iuwh 6mo ago[flagged]
- asim 6mo ago[dead]
- ggnore7452 6mo agotoo bad that only the smaller on-device models support native audio input.
- RandyOrion 6mo agoThank you Gemma team for releasing small dense VLM(s). The elo ranking [1] is too good to be true. I don't know why gemma-4-26b-a4b performs better than gemma-4-31b. Also waiting for more bugfixes in llama.cpp, sglang and vllm to do proper evaluations. [1] https://arena.ai/leaderboard/text/expert?license=open-source https://arena.ai/leaderboard/text/expert?license=open-source
- techpulselab 6mo ago[dead]
- EdoardoIaga 6mo ago[flagged]
- deleted 6mo ago[deleted]
- oblio 6mo agoHow do these compare to Open AI OSS?
- logicallee 6mo agoIf anyone here is interested in its creative writing style, I gave both the 10 GB and 20 GB models the prompt "write a short story", here the results: [1] They don't really have the structure of a short story, though the 20 GB model is more interesting and has two characters rather than just one character. In another comment, I gave them coding tasks, if you want to see how fast it does at coding (on a 24 GB Mac Mini M4 with 10 cores) you can watch me livestream this here: [2] Both models completed the fairly complex coding task well. [1] https://pastebin.com/ZcWv6Hkb https://pastebin.com/ZcWv6Hkb [2] https://www.youtube.com/live/G5OVcKO70ns https://www.youtube.com/live/G5OVcKO70ns
- Igor_Wiwi 6mo agoI created a blog post specifically about running these models locally on your machine (1 liner but getting gguf may take some time): https://igorstechnoclub.com/running-gemma-4-locally-in-almost-one-line/ https://igorstechnoclub.com/running-gemma-4-locally-in-almos...
- Praxwise 6mo agoI just checked the status of the domain registrations and noticed that the domain squatters have already started taking action. Almost all of the domains have been registered.
- lousken 6mo agoThe speed is complete poopoo, even on their API. To spend 5 seconds thinking about "hello how you doin" prompt on their TPUs is insane and something must be wrong with this model.
- zkmon 6mo agoIt would be helpful to know what kind of tasks does it beat Qwen models of similar size.
- konart 6mo agoSo many comments, but in the end - it can't write a simple set of unit tests in go with mockery.
- pratyushsood 6mo ago[dead]
- anonyfox 6mo agoM5 air here with 32gb ram and 10/10 cores. Anyone got some luck with mlx builds on oMLX so far? Not at my machine right now and would love to know if these models already work including tool calling
- aiiaro 6mo ago[flagged]
- kordlessagain 6mo agoIf you use Ollama: ollama pull gemma4:e2b # smallest ollama run gemma4:e2b # or larger: ollama pull gemma4:e4b ollama pull gemma4:26b ollama pull gemma4:31b
- mudkipdev 6mo agoIf you use the 'run' command, it pulls automatically for you
- mikewarot 6mo agoI updated Ollama (again) and changed my windows swap file settings to use up to 200 Gb of C: (an SSD). On the largest model (gemma4:31b), I seem to be getting about 5 tokens per second. This is amazing to me, because I'm using a $100 computer, without any fancy GPU. I love watching it "think". Consider this is thousands of times faster than any written conversations in the past. Those involved pieces of paper being transported, read, considered, replies written, then transported back. If it'll write code that doesn't completely suck, I think even this is good enough. What do you consider the lowest acceptable rate of generating tokens/second?
- mudkipdev 6mo agoUnder 15 is too slow for conversation personally. I guess 5 tokens per second is nice if you're one of the people who likes letting coding agents run overnight
- a96 6mo agoI think it depends on the kind of answer and how long the round trip is. Even a fast model that waffles for several minutes before giving an actual answer (coughqwen3.5cough) can feel very slow. Few seconds per word output may be fine if the final answer is correct and short. But generally, I'd like to see above 20, >50 is mostly great, and more is better. For conversational response, that is, not batch or interactive loop.
- bwannasek 6mo agoUsing Gemma 4 with OpenCode was more challenging than expected due to some active bugs in ollama related to reasoning and streaming - I did a quick writeup in how I used llama.cpp instead of ollama and how to set it up to support multi-turn tool calls properly in case this is helpful to others: https://bernhardwannasek.com/using-gemma-4-for-agentic-coding/ https://bernhardwannasek.com/using-gemma-4-for-agentic-codin...
- ronb1964 6mo agoI have Ollama installed on my Linux desktop with Alpaca as the frontend, but honestly I haven't done much with it beyond poking around. I also built a local speech-to-text app using Claude Code that runs Whisper offline, so I'm clearly drawn to the idea of keeping AI on-device. I'm curious whether Gemma 4 would be a noticeable step up for someone just using a local model for everyday tasks...writing, Q&A, that kind of thing. Is there a practical size recommendation for someone who isn't doing anything exotic, just wants a capable local model that doesn't require a supercomputer? And is there an advantage to having all this work with Claude somehow to broaden what is currently capable?
- gigatexal 6mo agoFor what it’s worth out the gate with ollama I can’t get it to work right in codex or claude. Seems to die after planning. Other models “just work” out of the box.
- lubitelpospat 6mo agoIf you're using litert-lm on a Mac with Apple Silicon - DO NOT forget to use "--backend gpu"! On my M1 Pro laptop this single setting resulted in 10x prefill performance and 2x decode performance. To anyone who knows how the internals of litert-lm work - what quantization does it use? How come the model is just 3.4 GB in size? EDIT: typo fix.
- hyperlambda 6mo ago[flagged]
- curioussquirrel 6mo agoFor anyone interested in multilingual performance, which is not usually well benchmarked or reported: Gemma 4 does really well, especially the dense 31B version. In fact, it outperforms many models with an order of magnitude higher number of parameters. It is not quite capable of performing work on really long tail languages, but their claim of 35 languages supported (and a hint of some knowledge of up to 140) was substantiated by our tests. If you're doing work outside of English and/or need to run a translation model in your terms, Gemma 4 is a very good candidate.
- om252345 6mo agoGemma 4 can unlock local agentic coding if coupled with right tools. I feel we need some graph based code kb, external memory and RAG will make it more powerful for local coding. I would say use Claude, gemini for big reactors but for small edits, using gemma 4 should be absolutely fine.