8 ms·
Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU
- neomindryan 3mo agoAuthor here. The short version: a viral post ran Gemma 4 on a 2016 Xeon; my Xeons are 2013, and the fork it used assumes AVX2, which Ivy Bridge doesn't have. The build failure was easy. The fun bug was the silent one: two MoE graph ops with no dispatch case on non-AVX2 builds, so every expert FFN output was uninitialized memory. Deterministic, NaN-free, fluent-looking multilingual gibberish. The fix is open upstream as PR #2138 (https://github.com/ikawrakow/ik_llama.cpp/pull/2138 https://github.com/ikawrakow/ik_llama.cpp/pull/2138), awaiting review. Fair warning on the AI angle: the patch was written by Claude at my direction. The post is explicit about which parts were me and which weren't. Happy to answer questions about either the bug or the workflow.
- otherjason 3mo agoThis reads as pretty clearly AI-generated text, which is against HN guidelines.
- pkghost 3mo agoHere's the thing: life also imitates art. If you invert your load-bearing assumption, it could be that he just reads too much slop. But my honest take? You might be right.
- dofm 3mo agostopitgetsomehelp.gif ;-)
- FL410 3mo agoThe PR? He said it was AI in the comment you replied to... I don't think the post itself reads like AI at all, but that's just me.
- logicallee 3mo agoI think "this" refers to its parent comment. Part of it sounds like Claude wrote it. AI-generated comments aren't allowed on HN.
- otherjason 3mo agoIndeed, I was referring to the parent comment.
- twoodfin 3mo agoThe post is absolutely LLM-generated. “Punchy” short sentences, “… has quietly come to mean …”, “The optimized paths weren’t there to execute.”
- why_only_15 3mo agoThe post itself is totally AI-generated. It has tons of tells, in addition to Pangram saying so. https://www.pangram.com/history/61fbc90e-180d-4d91-a85d-16ff5fc6ff87 https://www.pangram.com/history/61fbc90e-180d-4d91-a85d-16ff...
- why_only_15 3mo agoWhy is the post AI-generated? If you're going to make something for us to read you pay us the courtesy of actually writing it first. https://www.pangram.com/history/61fbc90e-180d-4d91-a85d-16ff5fc6ff87 https://www.pangram.com/history/61fbc90e-180d-4d91-a85d-16ff...
- neomindryan 3mo agoAuthor here, it looks like my original comment was flagged for some reason. The fix is open upstream as PR #2138 (https://github.com/ikawrakow/ik_llama.cpp/pull/2138 https://github.com/ikawrakow/ik_llama.cpp/pull/2138)
- deleted 3mo ago[deleted]
- aniwalunj 3mo agoTruly amazing. This gives a peek into the future for what's possible.
- throwaway2027 3mo agoThat's quite slow I'm getting 8-12 t/s on a 13 year old CPU. (Speed varies by context size and other settings who knows) https://news.ycombinator.com/item?id=48354801 https://news.ycombinator.com/item?id=48354801
- neomindryan 3mo agoThank you for sharing / linking!
- tuwtuwtuwtuw 3mo agoBut OP is using Q8 and you're using Q4?
- kQq9oHeAz6wLLS 3mo agoYeah, I'm seeing 8-9 t/s on a Xeon CPU E3-1270 V2 @ 3.50GHz with an old Nvidia Quadro K2200 (4GB). I run gemma4:e2b and gemma4:12b-it-qat on Ollama.
- hparadiz 3mo agoHere's my report running several different models on a dual Xeon with 256 GB of DDR4 and no GPU. https://gist.github.com/hparadiz/f3596d00a62d8ebb2dadcc46ee5822c7 https://gist.github.com/hparadiz/f3596d00a62d8ebb2dadcc46ee5...
- neomindryan 3mo agoThank you for sharing!
- puzzlingcaptcha 3mo agoHave you tried with a single CPU to get rid of the NUMA penalty? I understand this likely means halving the memory but I am interested in how much of a difference it makes
- trollbridge 3mo agoI have (192GB machine with two CPUs), pretty much does the trick. It just runs some small models used for embedding, etc. and has those on one CPU / memory node and all the Docker containers on the other one.c
- ryandrake 3mo agoI have a dual xeon also, same as OP: Ivy Bridge + 128GB DRAM, and was never really able to get decent LLM performance out of it. So I ended up biting the bullet and adding a "budget tier" A4000 20GB GPU. Too bad all my DRAM is wasted now--not sure if there is a way to take advantage of lots of DRAM once you move over to having inference happening on the GPU.
- puzzlingcaptcha 3mo agoHave you tried putting the KV cache on the GPU and running inference from RAM? From what I gather, prompt processing is particularly painful using RAM alone.
- ChrisArchitect 3mo agoRelated: A 10 year old Xeon is all you need https://news.ycombinator.com/item?id=48353348 https://news.ycombinator.com/item?id=48353348
- TacticalCoder 3mo agoYes and a 10 year old Xeon is going to be a v4 (not a v2 as in TFA) and it's going to have DDR4 ECC, not DDR3 ECC. I've got a 14 cores / 28 threads Xeon from 2015 that I use as a server at home (ZFS / VMs). It's really a sweet machine. For ricing I've got a semi-recent AMD 7700X / DDR5 RAM (from 2023 ?) which is my main machine but the real deal is my old and trusty 10 years old Xeon server. DDR4 ECC is pricey too atm but a 10 years old Xeon is basically free now. A 20 cores / 40 threads costs maybe 20 USD (for just the CPU). Slap that in a $100 old HP Z440 workstation and you're good to go for quite a few workloads. Mine is only on when I'm at my computer: it's not turned on 24/7 but more like 8/7 so the entire "but it consumes energy" point is moot.
- rvba 3mo agoApologies for asking here but literally nobody knows: Android studio connected to a local model disconnects automatically after 10 minutes. How set this limit to 12 hours or remove it completely? I could run my LM studio model all night... but I cant, since Android studio times out after a hard limit of 10M. This is not related to number of tokens. I tried Googling, searching for settings in Android studio, even created a stackoverflow post - but zero information. Jetbrains mentions "remote agent timeout mechanism" - but after changing it, nothing happens.
- NortySpock 3mo agoIf the local model is served via ollama, there's a default timeout of 10 minutes , which can be adjusted either per-call , or (as I did) in the systemd service environment variables https://docs.ollama.com/faq#how-do-i-keep-a-model-loaded-in-memory-or-make-it-unload-immediately https://docs.ollama.com/faq#how-do-i-keep-a-model-loaded-in-... You didn't specify what was serving your local model.
- rvba 3mo agoThank you for your reply. I use LM studio (local server), but can switch to a different tool. Do you know how to switch it in LM studio? What I see is that: android studio gives "Error: stream failed" and in LM studio server I see it is still working, then says that client (=android studio) disconnected. So I assumed it was a setting in android studio.
- NortySpock 3mo agoDunno, I have not used either of those. (Had been using zed and ollama, and ollama had plenty of odd defaults that needed fixing) Glancing through the docs, I would be digging down in the config of both Android studio and lm studio for either a TTL or jit auto evict setting, and if you find it, set it to some large number measured in hours? https://developer.android.com/studio/gemini/use-a-local-model https://developer.android.com/studio/gemini/use-a-local-mode... https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict
- mmastrac 3mo agoIs it just me or does this post not mention how much RAM they had? I would love to know - I have a dual-Xeon 1U screamer with 96GB of DDR4 RDIMM just sitting around... FWIW I'm getting a hardware max of 20 tok/s (approx topping out the GPU's compute) on my custom local diffusiongemma port running on an M3.
- neomindryan 3mo agohey, I’m the author. That box has 384gb, but loading the model “only” uses about 80gb.
- fouc 3mo agoany reason you went with q8 over q4? I'm wondering if q4 would run noticeably faster or not.
- superkuh 3mo agoSuch a system is RAM bandwidth limited and not compute limited Switching to q4 from q8 would decrease the amount of data needing to be loaded by half. The token generation rate would nearly double. But generally if you can do q6 or q8 and you have enough RAM you really should. Even if it's slower.
- giantrobot 3mo agoToken generation is nominally bandwidth limited. Prefill/prompt processing is nominally compute limited. For CPU inference on old hardware I don't think q4 offers any benefit over q8 since the AVX unit doesn't support such small floats. I don't even think AVX supports 4-bit int math. IIRC AVX2 does.
- neomindryan 3mo agoI think I was just following along with the previous post about running Gemma on a Xeon. Next I’m going to see which model can give the highest tokens/sec
- robotswantdata 3mo agoNeed to run this on my Xeons with AMX
- deltamidway 3mo agoHe's shown me his set up in his basement. It's sick! Talk about your 3d printer next!
- okokwhatever 3mo agoTo me context means everything. Tokens per second is a great metric but in the real world context window is the deal breaker when a real use case is on the table.
- dofm 3mo agoGemma 4 26B is capable up to 256k or 262k, can't remember which. Whether the writer's setup affects that choice I don't know.
- dwa3592 3mo agoI have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat https://github.com/deepanwadhwa/samosa-chat This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.
- embedding-shape 3mo ago> I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat https://github.com/deepanwadhwa/samosa-chat It's quite telling you didn't use Qwen3.6-35B-A3B locally to build that, seems there was another collaborator ;) Show something you've built with the model+tooling instead, truly dogfood it. I'm sure you'll discover things along the way too!
- dwa3592 3mo ago>>It's quite telling you didn't use Qwen3.6-35B-A3B locally to build that that would have run into a race condition unfortunately ;) but there is a sample landing page + a python function on the repo which shows what the model produced. my goal is to integrate the local model in my workflow so that claude/OAI can call this model for basic stuff.
- embedding-shape 3mo ago> that would have run into a race condition unfortunately ;) Not really, you start small, bootstrap as soon as you can, and off you go. Requires a good model though ;)
- smith7018 3mo agoNothing says they're using Qwen for local development. They could be using it to for conversations, knowledge, or "creative writing."
- 3mo ago
- OsamaJaber 3mo ago[dead]
- broabprobe 3mo agoI run the same setup Gemma 4 26B on a 2013 Mac Pro (dual graphics cards but they're useless for this). I also get about 5 t/s. It's perfectly serviceable for some tasks!
- kzrdude 3mo agoWhat is it useful for at slow speed?
- sjs382 3mo agoI bought a trashcan Mac Pro on a whim last week ($120 in eBay!) and did some reading about them—it turns out people recently started using the GPUs to run models @ 20-30 tok/s. I'm excited to get my mitts on it on Friday when it finally arrives. Here's some of the resources I came across if you're interested in reading. https://echalupa.com/blog/mac-pro-6-1-llama-cpp-firepro-d300-vulkan-ubuntu https://echalupa.com/blog/mac-pro-6-1-llama-cpp-firepro-d300... https://matthewgribben.com/blog/mac-pro-6-1-llama-cpp-firepro-d700-vulkan-ubuntu https://matthewgribben.com/blog/mac-pro-6-1-llama-cpp-firepr...
- broabprobe 3mo agooh very exciting! Thanks for sharing these sjs382! A bummer the models can't be run on both the GPUs and CPU so you could run much larger models. A year ago I was running gpt120b as it easily fit in 128gb ram. But now I'm only running gemma 26b and wishing I had stuck with 64gb ram since 128gb gets throttled. Not regretting buying 128gb 1.5 years ago though!
- rvz 3mo ago...and how many servings can this do?
- CurbStomper 3mo ago[dead]
- hagen8 3mo agoSome ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally. If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider. I estimate that the server consumes probably around 500W during inference. In Germany where 1kwh cost around 0.3USD, 18k tokens inferred locally would therefore cost 0.15USD which is 30x the costs of using an inference provider. But for ppl who worry about their data, running locally might still be good. However, they should be aware, that it is much less efficient than using an inference provider. The efficiency gap will also significantly increase as new GPUs will make inference much more efficient. EDIT: I first thought it'd be 180k token, but thanks to someone mentioning in the comments, it is 18k. I guess with that, it will be tough unless u got electricity almost for free. Also, the inference providers are probably still using H200/H100 for those small models. Once they use GB300 or next year the new Ruby GPUs, inference will be cheaper by a factor of 30. By then, running local models will mostly be about privacy.
- bigyabai 3mo agoIt's the "Race-to-Idle" situation all over again. It consumes less power to complete a task faster, whereas using "low power" hardware that draws max TDP for 30 minutes isn't very power efficient. The privacy nuts have a better leg to stand on, but even then it's hard to believe that they're using on-prem AI to replace SOTA model inference. As cool as local LLMs are, a lot of the stuff people run is a novelty.
- newusertoday 3mo agoits 18k not 180k
- sosodev 3mo agoOf course efficiency matters, but a lot of people either have cheap electricity or efficient hardware. My AMD strix halo home server can serve Gemma4-26B at like 70 TPS (rough estimate, I don’t remember the exact speed buts its fast af) while only using 100W.
- albrewer 3mo ago
- indigodaddy 3mo agoUnfortunately, the post comes off as AI-written. Why not just write your own posts?
- CookieCrisp 3mo agoNot everyone is as confident in their writing as they are in their engineering
- jeswin 3mo agoThe transformer architecture is fundamentally unsuitable for local inference, while being efficient at scale. It's a fun experiment to try, but that's about it.
- jeswin 3mo ago* autoregressive
- thomasjb 3mo agoI was inspired by that post also, got a Qwen Coder 1.5B up to 27tok/s prompt eval and 13tok/s decode on an e5-2650v2 inside a GNOME box
- rrhjm53270 3mo agoGood job. Thanks for sharing. I have a similar NAS server with 2 Intel(R) Xeon(R) CPU E5-2699 CPUs. I will test as well.
- Aurornis 3mo agoA dual Xeon of this era is probably pulling 300W or more when loaded. At national average electricity prices, that’s $1.35 per day. More during the summer if you have to cool the space. If you run it 24/7 and ignore prompt processing time (not a good assumption at all) it would get around 400,000 tokens in a day. That’s about $0.30 per million output tokens. Coincidentally, that’s the same price for this model on OpenRouter right now, but OpenRouter token gen will be 8X faster. There are a lot of good reasons to experiment with running LLMs locally, like if you don’t want any data leaving your house. Don’t think that you’re going to come out ahead monetarily. I say this as someone with a lot more money invested in local inference hardware at home. It’s fun, but it’s not a way to save money.
- brailsafe 3mo agoReasonable analysis, especially because this person seems to have an actual house. In my case, I rent and don't pay for electricity directly, so the cost effectiveness threshold is whenever the landlord starts complaining
- twsted 3mo agoI think, may be actually wrong, that most of us do not consider running a model locally a way to save money. It is a way not to spread personal info around.
- Aurornis 3mo agoAnyone running LLMs at home will come to that realization quickly, if they’re looking at their power bills. Even feeling the heat output of a computer running at 100% in your office makes it clear. I was responding to a lot of the comments saying this was a reasonable way to avoid paying for tokens or subscriptions. I don’t want anyone getting the wrong idea that this is a way to save money if that’s their priority.
- euio757 3mo ago> Even feeling the heat output of a computer running at 100% in your office makes it clear. What does it make clear? That I can replace the space heater my wife runs 9 out of 12 months of the year with a home server? And effectively get $0.00 per token during those times? In houses running A/C year round, sure there'd be some impact, but in all the places running heat, doesn't seem that it'd move the needle on power bills. There are startups whose entire business model is "cloud server as a home space heater" (aka "data furnace") ...
- rhema 3mo agoI love my little dual core X99 board with Xeon E5 2673 V3. It's not power efficient, but I just leave it in my basement for local Jupyter Notebook stuff. Much faster than everything cloud-based for a reasonably price at my scale.
- simonw 3mo agoHow much RAM did this need?
- 0x457 3mo ago[dead]
- throwawayffffas 3mo agoThe 5.2 tokens per second generation is not that bad, what kills it is the 16.2 prompt processing that makes this too slow to consider even if you have the hardware lying around.
- ColdStream 3mo agoVaguely related. Running an LLM on a Pentium 4. Nick named NetburstGPT. Yes, it is very slow! https://www.youtube.com/watch?v=ILV-eu90te8 https://www.youtube.com/watch?v=ILV-eu90te8
- linncharm 3mo ago[flagged]
- deleted 3mo ago[deleted]
- vugar82 3mo ago[dead]
- rbanffy 3mo agoI'm curious - these StoreVirtual machines don't seem to have any ports I could use to install my software on them. There is a USB port and that's about it. Is it installed with help of a serial console?
- ncgl 3mo agoIs more cores better here? 2690v2 vs vs 2673v2