15 ms·
Gemma 3 QAT Models: Bringing AI to Consumer GPUs
- jarbus 1y agoVery excited to see these kinds of techniques, I think getting a 30B level reasoning model usable on consumer hardware is going to be a game changer, especially if it uses less power.
- apples_oranges 1y agoDeepseek does reasoning on my home Linux pc but not sure how power hungry it is
- emrah 1y agoAvailable on ollama: https://ollama.com/library/gemma3 https://ollama.com/library/gemma3
- jinay 1y agoMake sure you're using the "-it-qat" suffixed models like "gemma3:27b-it-qat"
- Zambyte 1y agoHere are the direct links: https://ollama.com/library/gemma3:27b-it-qat https://ollama.com/library/gemma3:27b-it-qat https://ollama.com/library/gemma3:12b-it-qat https://ollama.com/library/gemma3:12b-it-qat https://ollama.com/library/gemma3:4b-it-qat https://ollama.com/library/gemma3:4b-it-qat https://ollama.com/library/gemma3:1b-it-qat https://ollama.com/library/gemma3:1b-it-qat
- ein0p 1y agoThanks. I was wondering why my open-webui said that I already had the model. I bet a lot of people are making the same mistake I did and downloading just the old, post-quantized 27B.
- Der_Einzige 1y agoHow many times do I have to say this? Ollama, llamacpp, and many other projects are slower than vLLM/sglang. vLLM is a much superior inference engine and is fully supported by the only LLM frontends that matter (sillytavern). The community getting obsessed with Ollama has done huge damage to the field, as it's ineffecient compared to vLLM. Many people can get far more tok/s than they think they could if only they knew the right tools.
- janderson215 1y agoI did not know this, so thank you. I read a blogpost a while back that encouraged using Ollama and never mention vLLM. Do you recommend reading any particular resource?
- 1y ago
- holografix 1y agoCould 16gb vram be enough for the 27b QAT version?
- halflings 1y agoThat's what the chart says yes. 14.1GB VRAM usage for the 27B model.
- erichocean 1y agoThat's the VRAM required just to load the model weights. To actually use a model, you need a context window. Realistically, you'll want a 20GB GPU or larger, depending on how many tokens you need.
- oezi 1y agoI didn't realize that the context would require such so much memory. Is this KV caches? It would seem like a big advantage if this memory requirement could be reduced.
- jffry 1y agoWith `ollama run gemma3:27b-it-qat "What is blue"`, GPU memory usage is just a hair over 20GB, so no, probably not without a nerfed context window
- woadwarrior01 1y agoIndeed, the default context length in ollama is a mere 2048 tokens.
- hskalin 1y agoWith ollama you could offload a few layers to cpu if they don't fit in the VRAM. This will cost some performance ofcourse but it's much better than the alternative (everything on cpu)
- 1y ago
- diggan 1y agoFirst graph is a comparison of the "Elo Score" while using "native" BF16 precision in various models, second graph is comparing VRAM usage between native BF16 precision and their QAT models, but since this method is about doing quantization while also maintaining quality, isn't the obvious graph of comparing the quality between BF16 and QAT missing? The text doesn't seem to talk about it either, yet it's basically the topic of the blog post.
- nithril 1y agoIn addition the graph "Massive VRAM Savings" graph states what looks like a tautology, reducing from 16 bits to 4 bits leads unsurprisingly to a x4 reduction in memory usage
- croemer 1y agoIndeed, the one thing I was looking for was Elo/performance of the quantized models, not how good the base model is. Showing how much memory is saved by quantization in a figure is a bit of an insult to the intelligence of the reader.
- claiir 1y agoYea they mention a “perplexity drop” relative to naive quantization, but that’s meaningless to me. > We reduce the perplexity drop by 54% (using llama.cpp perplexity evaluation) when quantizing down to Q4_0. Wish they showed benchmarks / added quantized versions to the arena! :>
- wtcactus 1y agoThey keep mentioning the RTX 3090 (with 24 GB VRAM), but the model is only 14.1 GB. Shouldn’t it fit a 5060 Ti 16GB, for instance?
- oktoberpaard 1y agoWith a 128K context length and 8 bit KV cache, the 27b model occupies 22 GiB on my system. With a smaller context length you should be able to fit it on a 16 GiB GPU.
- jsnell 1y agoMemory is needed for more than just the parameters, e.g. the KV cache.
- cubefox 1y agoKV = key-value
- Havoc 1y agoJust checked - 19 gigs with 8k context @ q8 kv.Plus another 2.5-ish or so for OS etc. ...so yeah 3090
- noodletheworld 1y ago? Am I missing something? These have been out for a while; if you follow the HF link you can see, for example, the 27b quant has been downloaded from HF 64,000 times over the last 10 days. Is there something more to this, or is just a follow up blog post? (is it just that ollama finally has partial (no images right?) support? Or something else?)
- deepsquirrelnet 1y agoQAT “quantization aware training” means they had it quantized to 4 bits during training rather than after training in full or half precision. It’s supposedly a higher quality, but unfortunately they don’t show any comparisons between QAT and post-training quantization.
- noodletheworld 1y agoI understand that, but the qat models (1) are not new uploads. How is this more significant now than when they were uploaded 2 weeks ago? Are we expecting new models? I don’t understand the timing. This post feels like it’s two weeks late. [1] - https://huggingface.co/collections/google/gemma-3-qat-67ee61ccacbf2be4195c265b https://huggingface.co/collections/google/gemma-3-qat-67ee61...
- llmguy 1y ago8 days is closer to 1 week then 2. And it’s a blog post, nobody owes you realtime updates.
- noodletheworld 1y agohttps://huggingface.co/google/gemma-3-27b-it-qat-q4_0-gguf/tree/main https://huggingface.co/google/gemma-3-27b-it-qat-q4_0-gguf/t... > 17 days ago Anywaaay... I'm literally asking, quite honestly, if this is just an 'after the fact' update literally weeks later, that they uploaded a bunch of models, or if there is something more significant about this I'm missing.
- behnamoh 1y agoThis is what local LLMs need—being treated like first-class citizens by the companies that make them. That said, the first graph is misleading about the number of H100s required to run DeepSeek r1 at FP16. The model is FP8.
- freeamz 1y agoso what is the real comparison against DeepSeek r1 ? Would be good to know which is actually more cost efficient and open (reproducible build) to run locally.
- behnamoh 1y agohalf the amount of those dots is what it takes. but also, why compare a 27B model with a +600B? that doesn't make sense.
- smallerize 1y agoIt's an older image that they just reused for the blog post. It's on https://ai.google.dev/gemma https://ai.google.dev/gemma for example
- mmoskal 1y agoAlso ~noone runs h100 at home, ie at batch size 1. What matters is throughput. With 37b active parameters and a massive deployment throughout (per gpu) should be similar to Gemma.
- mythz 1y agoThe speed gains are real, after downloading latest QAT gemma3:27b eval perf is now 1.47x faster on ollama, up from 13.72 to 20.11 tok/s (on A4000's).
- btbuildem 1y agoIs 27B the largest QAT Gemma 3? Given these size reductions, it would be amazing to have the 70B!
- umajho 1y agoI am currently using the Q4_K_M quantized version of gemma-3-27b-it locally. I previously assumed that a 27B model with image input support wouldn't be very high quality, but after actually using it, the generated responses feel better than those from my previously used DeepSeek-R1-Distill-Qwen-32B (Q4_K_M), and its recognition of images is also stronger than I expected. (I thought the model could only roughly understand the concepts in the image, but I didn't expect it to be able to recognize text within the image.) Since this article publishes the optimized Q4 quantized version, it would be great if it included more comparisons between the new version and my currently used unoptimized Q4 version (such as benchmark scores). (I deliberately wrote this reply in Chinese and had gemma-3-27b-it Q4_K_M translate it into English.)
- rob_c 1y agoGiven how long between this being released and this community picking up on it... Lol
- GaunterODimm 1y ago2days :/...
- rob_c 1y agoGiven I know people running gemma3 on local devices for over almost a month now this is either a very slow news day or evidence of finger missing the pulse... https://blog.google/technology/developers/gemma-3/ https://blog.google/technology/developers/gemma-3/
- simonw 1y agoThis is new. These are new QAT (Quantization-Aware Training) models released by the Gemma team.
- rob_c 1y agoThere's nothing more than an iteration on the topic, gemma3 was smashing local results a month ago and made no waves as it dropped...
- simonw 1y agoQuoting the linked story: > Last month, we launched Gemma 3, our latest generation of open models. Delivering state-of-the-art performance, Gemma 3 quickly established itself as a leading model capable of running on a single high-end GPU like the NVIDIA H100 using its native BFloat16 (BF16) precision. > To make Gemma 3 even more accessible, we are announcing new versions optimized with Quantization-Aware Training (QAT) that dramatically reduces memory requirements while maintaining high quality. The thing that's new, and that is clearly resonating with people, is the "To make Gemma 3 even more accessible..." bit.
- simonw 1y agoI think gemma-3-27b-it-qat-4bit is my new favorite local model - or at least it's right up there with Mistral Small 3.1 24B. I've been trying it on an M2 64GB via both Ollama and MLX. It's very, very good, and it only uses ~22Gb (via Ollama) or ~15GB (MLX) leaving plenty of memory for running other apps. Some notes here: https://simonwillison.net/2025/Apr/19/gemma-3-qat-models/ https://simonwillison.net/2025/Apr/19/gemma-3-qat-models/ Last night I had it write me a complete plugin for my LLM tool like this: llm install llm-mlx llm mlx download-model mlx-community/gemma-3-27b-it-qat-4bit llm -m mlx-community/gemma-3-27b-it-qat-4bit \ -f https://raw.githubusercontent.com/simonw/llm-hacker-news/refs/heads/main/llm_hacker_news.py \ -f https://raw.githubusercontent.com/simonw/tools/refs/heads/main/github-issue-to-markdown.html \ -s 'Write a new fragments plugin in Python that registers issue:org/repo/123 which fetches that issue number from the specified github repo and uses the same markdown logic as the HTML page to turn that into a fragment' It gave a solid response! https://gist.github.com/simonw/feccff6ce3254556b848c27333f52543#response https://gist.github.com/simonw/feccff6ce3254556b848c27333f52... - more notes here: https://simonwillison.net/2025/Apr/20/llm-fragments-github/ https://simonwillison.net/2025/Apr/20/llm-fragments-github/
- littlestymaar 1y ago> and it only uses ~22Gb (via Ollama) or ~15GB (MLX) Why is the memory use different? Are you using different context size in both set-ups?
- simonw 1y agoNo idea. MLX is its own thing, optimized for Apple Silicon. Ollama uses GGUFs. https://ollama.com/library/gemma3:27b-it-qat https://ollama.com/library/gemma3:27b-it-qat says it's Q4_0. https://huggingface.co/mlx-community/gemma-3-27b-it-qat-4bit https://huggingface.co/mlx-community/gemma-3-27b-it-qat-4bit says it's 4bit. I think those are the same quantization?
- jychang 1y agoThose are the same quant, but this is a good example of why you shouldn't use ollama. Either directly use llama.cpp, or use something like LM Studio if you want something with a GUI/easier user experience. The Gemma 3 17b QAT GGUF should be taking up ~15gb, not 22gb.
- justanotheratom 1y agoAnyone packaged one of these in an iPhone App? I am sure it is doable, but I am curious what tokens/sec is possible these days. I would love to ship "private" AI Apps if we can get reasonable tokens/sec.
- Alifatisk 1y agoIf you ever ship a private AI app, don't forget to implement the export functionality, please!
- deleted 1y ago[deleted]
- idonotknowwhy 1y agoYou mean conversations? Just the jsonl of the standard hf dataset format to import into other systems?
- Alifatisk 1y agoYeah I mean conversations.
- nico 1y agoWhat kind of functionality do you need from the model? For basic conversation and RAG, you can use tinyllama or qwen-2.5-0.5b, both of which run on a raspberry pi at around 5-20 tokens per second
- justanotheratom 1y agoI am looking for structured output at about 100-200 tokens/second on iPhone 14+. Any pointers?
- nico 1y agoThe qwq-2.5-0.5b is the tiniest useful model I've used, and pretty easy to fine-tune locally on a Mac. Haven't tried it on an iPhone, but given it runs at about 150-200 tokens/second on a Mac, I'm kinda doubtful it could do the same on an iPhone. But I guess you'd just have to try
- Alifatisk 1y agoExcept this being lighter than the other models, is there anything else the Gemma model is specifically good at or better than the other models at doing?
- itake 1y agoGoogle claims to have better multi language support, due tokenizer improvements.
- nico 1y agoThey are multimodal. Havent tried the QAT one yet. But the gemma3s released a few weeks ago are pretty good at processing images and telling you details about what’s in them
- Zambyte 1y agoI have found Gemma models are able to produce useful information about more niche subjects that other models like Mistral Small cannot, at the expense of never really saying "I don't know", where other models will, and will instead produce false information. For example, if I ask mistral small who I am by name, it will say there is no known notable figure by that name before the knowledge cutoff. Gemma 3 will say I am a well known <random profession> and make up facts. On the other hand, I have asked both about local organization in my area that I am involved with, and Gemma 3 could produce useful and factual information, where Mistral Small said it did not know.
- trebligdivad 1y agoIt seems pretty impressive - I'm running it on my CPU (16 core AMD 3950x) and it's very very impressive at translation, and the image description is very impressive as well. I'm getting about 2.3token/s on it (compared to under 1/s on the Calme-3.2 I was previously using). It does tend to be a bit chatty unless you tell it not to be; pretty much everything it'll give you a 'breakdown' unless you tell it not to - so for traslation my prompt is 'Translate the input to English, only output the translation' to stop it giving a breakdown of the input language.
- simonw 1y agoWhat are you using to run it? I haven't got image input working yet myself.
- trebligdivad 1y agoI'm using llama.cpp - built last night from head; to do image stuff you have to run a separate client they provide, with something like: ./build/bin/llama-gemma3-cli -m /discs/fast/ai/gemma-3-27b-it-q4_0.gguf --mmproj /discs/fast/ai/mmproj-model-f16-27B.gguf -p "Describe this image." --image ~/Downloads/surprise.png Note the 2nd gguf in there - I'm not sure, but I think that's for encoding the image.
- terhechte 1y agoImage input has been working with LM Studio for quite some time
- tough 1y agoneed it here for cli usage! https://github.com/agustif/llm-lmstudio https://github.com/agustif/llm-lmstudio
- Havoc 1y agoThe upcoming qwen3 series is supposed to be MoE...likely to give better tk/s on CPU
- XCSme 1y agoSo how does 27b-it-qat (18GB) compare to 27b-it-q4_K_M (17GB)?
- perching_aix 1y agoThis is my first time trying to locally host a model - gave both the 12B and 27B QAT models a shot. I was both impressed and disappointed. Setup was piss easy, and the models are great conversationalists. I have a 12 gig card available and the 12B model ran very nice and swift. However, they're seemingly terrible at actually assisting with stuff. Tried something very basic: asked for a powershell one liner to get the native blocksize of my disks. Ended up hallucinating fields, then telling me to go off into the deep end, first elevating to admin, then using WMI, then bringing up IOCTL. Pretty unfortunate. Not sure I'll be able to put it to actual meaningful use as a result.
- parched99 1y agoI think Powershell is a bad test. I've noticed all local models have trouble providing accurate responses to Powershell-related prompts. Strangely, even Microsoft's model, Phi 4, is bad at answering these questions without careful prompting. Though, MS can't even provide accurate PS docs. My best guess is that there's not enough discussion/development related to Powershell in training data.
- fragmede 1y agoWhich, like, you'd think Microsoft has an entire team there who's purpose would be to generate good PowerShell for it to train on.
- terhechte 1y agoLocal models, due to their size more than big cloud models, favor popular languages rather than more niche ones. They work fantastic for JavaScript, Python, Bash but much worse at less popular things like Clojure, Nim or Haskell. Powershell is probably on the less popular side compared to Js or Bash. If this is your main use case you can always try to fine tune a model. I maintain a small llm bench of different programming languages and the performance difference between say Python and Rust on some smaller models is up to 70%
- perching_aix 1y ago
- CyberShadow 1y agoHow does it compare to CodeGemma for programming tasks?
- api 1y agoWhen I see 32B or 70B models performing similarly to 200+B models, I don’t know what to make of this. Either the latter contains more breadth of information but we have managed to distill latent capabilities to be similar, the larger models are just less efficient, or the tests are not very good.
- simonw 1y agoIt makes intuitive sense to me that this would be possible, because LLMs are still mostly opaque black boxes. I expect you could drop a whole hunch of the weights without having a huge impact on quality - maybe you end up mostly ditching the parts that are derived from shitposts on Reddit but keep the bits from Arxiv for example. (That's a massive simplification of how any of this works, but it's how I think about it at a high level.)
- retinaros 1y agoits just bs benchmarks. they are all cheating at this point feeding the data in the training set. doesnt mean the llm arent becoming better but when they all lie...
- mekpro 1y agoGemma 3 is way way better than Llama 4. I think Meta will start to lose its position in LLM mindshare. Another weakness of Llama 4 is its model size that is too large (even though it can run fast with MoE), which greatly limits the applicable users to a small percentage of enthusiasts who have enough GPU VRAM. Meanwhile, Gemma 3 is widely usable across all hardware sizes.
- miki123211 1y agoWhat would be the best way to deploy this if you're maximizing for GPU utilization in a multi-user (API) scenario? Structured output support would be a big plus. We're working with a GPU-poor organization with very strict data residency requirements, and these models might be exactly what we need. I would normally say VLLM, but the blog post notably does not mention VLLM support.
- PhilippGille 1y agovLLM lists Gemma 3 as supported, if I'm not mistaken: https://docs.vllm.ai/en/latest/models/supported_models.html#list-of-text-only-language-models https://docs.vllm.ai/en/latest/models/supported_models.html#...
- 999900000999 1y agoAssuming this can match Claude's latest, and full time usage ( as in you have a system that's constantly running code without any user input,) you'd probably save 600 to 700 a month. A 4090 is only 2K and you'll see an ROI within 90 days. I can imagine this will serve to drive prices for hosted llms lower. At this level any company that produces even a nominal amount of code should be running LMS on prem( AWS if your on the cloud).
- rafaelmn 1y agoI'd say using a Mac studio with M4 Max and 128 GB RAM will get you way further than 4090 in context size and model size. Cheaper than 2x4090 and less power while being a great overall machine. I think these consumer GPUs are way too expensive for the amount of memory they pack - and that's intentional price discrimination. Also the builds are gimmicky. It's just not setup for AI models, and the versions that are cost 20k. AMD has that 128GB RAM strix halo chip but even with soldered ram the bandwidth there is very limited, half of M4 Max, which is half of 4090. I think this generation of hardware and local models is not there yet - would wait for M5/M6 release.
- tootie 1y agoThere's certainly room to grow but I'm running Gemma 12b on a 4060 (8GB VRAM) which I bought for gaming and it's a tad slow but still gives excellent results. And it certainly seems software is outpacing hardware right now. The target is making a good enough model that can run on a phone.
- retinaros 1y agotwo 3090 are the way to go
- briandear 1y agoThe normal Gemma models seem to work fine on Apple silicon with Metal. Am I missing something?
- simonw 1y agoThese new special editions of those models claim to work better with less memory.
- porphyra 1y agoIt is funny that Microsoft had been peddling "AI PCs" and Apple had been peddling "made for Apple Intelligence" for a while now, when in fact usable models for consumer GPUs are only barely starting to be a thing on extremely high end GPUs like the 3090.
- ivape 1y agoThis is why the "AI hardware cycle is hype" crowd is so wrong. We're not even close, we're basically at ColecoVision/Atari stage of hardware here. It's going be quite a thing when everyone gets a SNES/Genesis.
- icedrift 1y agoCapable local models have been usable on Macs for a while now thanks to their unified memory.
- NorwegianDude 1y agoA 3090 is not a extremely high end GPU. Is a consumer GPU launched in 2020, and even in price and compute it's around a mid-range consumer GPU these days. The high end consumer card from Nvidia is the RTX 5090, and the professional version of the card is the RTX PRO 6000.
- zapnuk 1y agoA 3090 still costs 1800€. Thats not mid-range by a long shot The 5070 or 5070ti are mid range. They cost 650/900€.
- NorwegianDude 1y ago3090s are no longer produced, that's why new ones are so expensive. At least here, used 3090s are around €650, and a RTX 5070 is around €625. It's definitely not extremely high end any more, the price is(at least here) the same as the new mid range consumer cards. I guess the price can vary by location, but €1800 for a 3090 is crazy, that's more than the new price in 2020.
- 1y ago
- mark_l_watson 1y agoIndeed!! I have swapped out qwen2.5 for gemma3:27b-it-qat using Ollama for routine work on my 32G memory Mac. gemma3:27b-it-qat with open-codex, running locally, is just amazingly useful, not only for Python dev, but for Haskell and Common Lisp also. I still like Gemini 2.5 Pro and o3 for brainstorming or working on difficult problems, but for routine work it (simply) makes me feel good to have everything open source/weights running on my own system. Wen I bought my 32G Mac a year ago, I didn't expect to be so happy as running gemma3:27b-it-qat with open-codex locally.
- Tsarp 1y agoWhat tps are you hitting? And did you have to change KV size?
- nxobject 1y agoFellow owner of a 32GB MBP here: how much memory does it use while resident - or, if swapping happens, do you see the effects in your day to day work? I’m in the awkward position of using on a daily basis a lot of virtualized bloated Windows software (mostly SAS).
- mark_l_watson 1y agoI have the usual programs running on my Mac, along with open-codex: Emacs, web browser, terminals, VSCode, etc. Even with large contexts, open-codex with Ollama and Gemma 3 27B QAT does not seem to overload my system. To be clear, I sometimes toggle open-codex to use the Gemini 3.5 Pro API also, but I enjoy running locally for simpler routine work.
- pantulis 1y agoHow did you manage to run open-codex against a local ollama? I keep getting 400 Errors no matter what I try with the --provider and --model options.
- pantulis 1y agoNever mind, found your Leanpub book and followed the instructions and at least I have it running with qwen-2.5. I'll investigate what happens with Gemma.
- piyh 1y agoMeta Maverick is crying in the shower getting so handily beat by a model with 15x fewer params
- Samin100 1y agoI have a few private “vibe check” questions and the 4 bit QAT 27B model got them all correctly. I’m kind of shocked at the information density locked in just 13 GB of weights. If anyone at Deepmind is reading this — Gemma 3 27B is the single most impressive open source model I have ever used. Well done!
- itake 1y agoI tried to use the -it models for translation, but it completely failed at translating adult content. I think this means I either have to train the -pt model with my own instruction tuning or use another provider :(
- andhuman 1y agoHave you tried Mistral Small 24b?
- itake 1y agoMy current architecture is an on-device model for fast translation and then replace that with a slow translation (via an API call) when its ready. 24b would be too small to run on device and I'm trying to keep my cloud costs low (meaning I can't afford to host a small 24b 24/7).
- jychang 1y agoTry mradermacher/amoral-gemma3-27B-v2-qat-GGUF
- itake 1y agoMy current architecture is an on-device model for fast translation and then replace that with a slow translation (via an API call) when its ready. 24b would be too small to run on device and I'm trying to keep my cloud costs low (meaning I can't afford to host a small 27b 24/7).
- cheriot 1y agoIs there already a Helium for GPUs?
- punnerud 1y agoJust tested the 27B, and it’s not very good at following instructions and is very limited on more complex code problems. Mapping from one JSON with a lot of plain text, into a new structure and it fails every time. Ask it to generate SVG, and it’s very simple and almost too dumb. Nice that it doesn’t need that huge amount of RAM, and perform ok on smaller languages from my initial tests.
- deleted 1y ago[deleted]
- Havoc 1y agoDefinitely my current fav. Also interesting that for many questions the response is very similar to the gemini series. Must be sharing training datasets pretty directly.
- mattfrommars 1y agoanyone had success using Gemma 3 QAT models on Ollama with cline? They just don't work as good compared Gemini 2.0 flash provided by API
- technologesus 1y agoJust for fun I created a new personal benchmark for vision-enabled LLMs: playing minecraft. I used JSON structured output in LM Studio to create basic controls for the game. Unfortunately no matter how hard I proompted, gemma-3-27b QAT is not really able to understand simple minecraft scenarios. It would say things like "I'm now looking at a stone block. I need to break it" when it is looking out at the horizon in the desert. Here is the JSON schema: https://pastebin.com/SiEJ6LEz https://pastebin.com/SiEJ6LEz System prompt: https://pastebin.com/R68QkfQu https://pastebin.com/R68QkfQu
- jvictor118 1y agoi've found the vision capabilities are very bad with spatial awareness/reasoning. They seem to know that certain things are in the image, but not where they are relative to each other, their relative sizes, etc.
- gigel82 1y agoFWIW, the 27b Q4_K_M takes about 23Gb of VRAM with 4k context and 29Gb with 16k context and runs at ~61t/s on my 5090.
- manjunaths 1y agoI am running this on 16 GB AMD Radeon 7900 GRE with 64 GB machine with ROCm and llama.cpp on Windows 11. I can use Open-webui or the native gui for the interface. It is made available via an internal IP to all members of my home. It runs at around 26 tokens/sec and FP16, FP8 is not supported by the Radeon 7900 GRE. I just love it. For coding QwQ 32b is still king. But with a 16GB VRAM card it gives me ~3 tokens/sec, which is unusable. I tried to make Gemma 3 write a powershell script with Terminal gui interface and it ran into dead-ends and finally gave up. QwQ 32B performed a lot better. But for most general purposes it is great. My kid's been using it to feed his school textbooks and ask it questions. It is better than anything else currently. Somehow it is more "uptight" than llama or the chinese models like Qwen. Can't put my finger on it, the Chinese models seem nicer and more talkative.
- mdp2021 1y ago> My kid's been using it to feed his school textbooks and ask it questions Which method are you employing to feed a textbook into the model?
- anshumankmr 1y agomy trusty RTX 3060 is gonna have its day in the sun... though I have run a bunch of 7B models fairly easily on Ollama.
- ece 1y agoOn Hugging Face: https://huggingface.co/collections/google/gemma-3-qat-67ee61ccacbf2be4195c265b https://huggingface.co/collections/google/gemma-3-qat-67ee61...
- yuweiloopy2 1y agoBeen using the 27B QAT model for batch processing 50K+ internal documents. The 128K context is game-changing for our legal review pipeline. Though I wish the token generation was faster - at 20tps it's still too slow for interactive use compared to Claude Opus.
- gitroom 1y agonice, loving the push with local models lately - always makes me wonder though, you think privacy wins out over speed and convenience in the long run or people just stick with what's quickest?
- simonw 1y agoSpeed and convenience will definitely win for most people. Hosted LLMs are so cheap these days, and are massively more capable than anything you can fit on even a very beefy ($4,000+) consumer machine. The privacy concerns are honestly mostly imaginary at this point, too. Plenty of hosted LLM vendors will promise not to train on your data. The bigger threat is if they themselves log data and then have a security incident, but honestly the risk that your own personal machine gets stolen or hacked is a lot higher than that.
- casey2 1y agoI don't get the appeal. For LLMs to be useful at all you at least need to bin the the dozen exabit range per token, zettabit/s if you want something usable. There is really no technological path towards supercomputers that fast in a human timescale and in 100 years. The thing that makes LLMs usefull is their ability to translate concepts from one domain to the other. Overfitting on choice benchmarks, even a spread, will lower their usefullness in every general task by destorying infomation that is encoded in the weights. Ask gemma to write a 5 paragraph essay on any niche topic and you will get plenty of statements that have an extremely small likely of existing in relation to the topic, but have a high likely of existing in related more popular topics. ChatGPT less so, but still at least one a paragraph. I'm not talking about factual errors or common oversimplifications. I'm talking about completely unrelated statements. What your asking about is largely outside it's training data of which a 27GB models gives you what? a few hundred Gigs? Seems like alot, but you have to remember that there is a lot of stuff that you probably don't care about that many people do. Stainless steel and Kubernetes are going to be well represented, your favorite media? probably not, relatively current? definitely not. Which sounds fine, until you realize that people who care about Stainless steel and Kubernetes, likely care about some much more specific aspect which isn't going to be represented and you are back to the same problem of low usability. This is why I believe that scale is king and that both data and compute are the big walls. Google has Youtube data but they are only using it in Gemini.
- deleted 1y ago[deleted]