8 ms·
I don't know about good, I use a lot of local models and they're still pretty painful to run locally You have dense models (qwen 27b, gemma 31b) who are pretty
by c0rruptbytes 4mo ago
I don't know about good, I use a lot of local models and they're still pretty painful to run locally
You have dense models (qwen 27b, gemma 31b) who are pretty smart, but pretty slow
You have MoE models (gemma 26b, qwen 35b, north mini code 30b) who are pretty fast, but make a lot of mistakes
You need a lot of memory to run these well, quantization makes tool calling weaker, so most run at 4 bit quants and are wondering why it kinda sucks and that's because you've essentially lobotomized the model (I recommend unsloth quants, i recommend 6bit for MoEs and 5bit for dense)
So you need a lot of compute to make the pre-fill fast, you need bandwidth to make the decode fast, you need a lot of memory to hold everything - lot of ifs
On top of that, your laptop becomes a loud hot churning machine, it's uncomfortable to work with.
So are they good? not really. Do they work? yes
edit: just wanna clarify - i think open models are the future, i think they're super important, i'm contributing constantly to the ecosystem - i think people should play around with these models, i think people should use `pi` and learn how it all works - but don't download a model expecting it to be good out of the box, you will have to tune and configure a lot of stuff to replace a "coding agent" that most people are using models for
- heipei 4mo agoDepends on what you mean by "local". On your Macbook, large dense models like Qwen 3.6 27B will be slow, sure. On a local workstation with a dedicated RTX card you can get > 100 tps, which is more than good enough to work with it, and faster than cloud models in many cases.
- jstanley 4mo agoBut how smart is it? All the people running local models never seem to mention that they are way dumber than cloud models. I don't care how many tokens per second of nonsense it can generate.
- myaccountonhn 4mo agoIts not going to be as good as Claude, but if you know what you're doing, it may be good enough to get your work done.
- data-ottawa 4mo agoThis is task dependent. I find devstral (even though it’s weak generally) much better at writing and documentation than Opus. I’m actually now delegating all documentation to devstral and away from Claude, which makes a mess.
- garciasn 4mo agoA highly skilled carpenter may be able to 'get work done' by banging nails in with a heavy-bottomed cocktail glass, doesn't mean it's not painful to do so when it is continuously breaking and leaving shards of glass all over the workshop for you to find every day for the rest of your life until you clean up the mess you made using the wrong tool for the job.
- CamperBob2 4mo agoMore like, a highly-skilled carpenter can work miracles with a $6 hammer from the hardware store, while the pros on the commercial crew are using fancy compressed-air tools. The carpenter has to get up close and personal with the wood. He can't match the crew's throughput, but maybe that's not what he's trying to do.
- mcbits 4mo agoI would say the hammer is no AI. Local models are the cheapest XKGYAGH electric nailer on Amazon that "works" but jams up all the time. The $20/mo cloud models are a nice DeWalt that gives an hour of jam-free operation but takes five hours to recharge. And if someone else is paying for it, one can use the heavy duty nail gun with a big generator and compressor on a trailer that can run all day.
- sgt101 4mo agoIf someone comes into the workshop and takes all the tools (hello Donald) then having a cocktail glass to hand might be a bit of a lucky break. (geddit?)
- heipei 4mo agoIt is smart enough that I use for all my coding tasks, and a lot of other mundane tasks. It is probably not smart enough for "design this whole architecture of this complex system from scratch, make no mistakes", but that is not something I want from a coding tool anyway. I want a model that I can point to a file and tell it to make some changes to the file and related files. Or that I can ask to review a PR with regards to certain aspects. My suggestion is to simply try it and see what it feels like.
- notnullorvoid 4mo agoQuantized Gemma 4 26B is as smart or better than GPT 5 in most of my testing. Granted GPT 5 is nearly a year old at this point, but I can run Gemma 4 on a ~6 year old consumer GPU (RTX 3090) and get 140 t/s.
- throwawayffffas 4mo agoQwen 3.6 35b a3b is about as good as sonnet 4.5. It varies but it's at that level.
- int_19h 4mo agoNot even close. It may be "about as good" on some very specific task.
- lelanthran 4mo ago> But how smart is it? All the people running local models never seem to mention that they are way dumber than cloud models. Well, you aren't going to give it a 20k line sec and have it churn out a full app after 4 hours hours. But, you can get it to write code for you if you do the design.
- c0rruptbytes 4mo agoI'm talking about the common use case that I think hacker news people have: you get a macbook for work, you run the macbook they're not going to start giving GPUs to employees to run local models
- zozbot234 4mo agoMaybe we shouldn't be running these models on laptops with their thermally constrained form factor, and we shouldn't expect quick inference on a par with a large cloud-based platform either, at least not for near-SOTA model quality. It's still worth it to avoid becoming massively reliant on centralized services.
- greenavocado 4mo agoI have a 5070 12 GB laptop GPU and can hit 72 tokens per second in the first couple thousand tokens before dropping to mid-high 50s after about 15k context. This setup is extremely optimized down to the last flag. Changing any param above the temp flag craters performance. I don't have enough system RAM to properly handle the large context windows so I don't use local models. # 1,257 tokens 17s 72.18 t/s $env:CUDA_DEVICE_SCHEDULE = "SPIN" cd D:\src\llama.cpp\ .\build\bin\Release\llama-server.exe ` --port 8080 ` --host 127.0.0.1 ` -m "D:\LLM\Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf" ` -fitt 2048 ` -c 98304 ` -n 32768 ` -fa on ` -np 1 ` --kv-unified ` -ctk q8_0 ` -ctv q8_0 ` -ctkd q8_0 ` -ctvd q8_0 ` -ctxcp 64 ` --mlock ` --no-warmup ` --spec-type draft-mtp ` --spec-draft-n-max 2 ` --spec-draft-p-min 0.1 ` --chat-template-kwargs '{\"preserve_thinking\": true}' ` --temp 0.6 ` --top-p 0.95 ` --top-k 20 ` --min-p 0.0 ` --presence-penalty 0.0 ` --repeat-penalty 1.0
- mattmanser 4mo agoThat's a quant 4 which the thread OP specifically called out as rubbish. The Q4_K_XL bit for those not in the know.
- greenavocado 4mo agoI typically find myself using a context of between 150-500k with GPT models so local models are simply not enough and I stopped using them.
- stymaar 4mo ago
- greenavocado 4mo ago4 bit unsloth quants are good if you never ask for more than 20k context, use it as autocomplete on steroids, and never delegate serious questions to it
- adam_arthur 4mo agoGemma 4 is particularly good at pipeline/automation tasks. It outperforms all the Qwen models (even 100B+) for rule following/automation style tasks in my experience. Its image interpretation is also very good, and out-benchmarks Opus. Qwen seems to ignore instructions and consistently outputs incorrect formats (when token generation format is not explicitly constrained) But yes, on the DGX Spark Gemma 31B Q4 with MTP runs around 20 tok/s and Gemma 26B A4B around 60 tok/s. Still quite slow. But on a high end Nvidia card would run significantly faster and still fit in memory. I'd recommend for anyone getting into local models to focus on memory bandwidth over RAM. Models under 100B parameters are now sufficient and hugely useful for automation. I agree that for coding/creation use cases, there's still not a compelling argument for local models. But e.g. if you want to scan a list of stocks and interpret news/high pass filtering, interpreting logs, interpreting screenshots, the local models are more than sufficient already.
- trouve_search 4mo agoOn a 5090, gemma4 26B runs at 350TPS with the command below [1] and gemma4 31B is around 150TPS with a similar command. I'm really surprised how much slower a DGX spark is for the same price. 1. Here's my command. PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ vllm serve cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit \ --dtype auto \ --gpu-memory-utilization 0.95 \ --kv-cache-dtype fp8 \ --enable-chunked-prefill \ --enable-prefix-caching \ --trust-remote-code \ --enable-auto-tool-choice \ --tool-call-parser gemma4 \ --reasoning-parser gemma4 \ --max-num-batched 16000 \ --max-model-len 64000 \ --max-num-seqs 12 --speculative-config '{"model": "./gemma-4-26B-A4B-it-assistant", "num_speculative_tokens": 4}'
- adam_arthur 4mo agoYes, I'd recommend a 5090 over the DGX Spark if your goal is general automation. You can run multiple instances of these models in parallel on the DGX Spark which somewhat mitigates the difference if your task is parallelizable. But I'd take the simplicity of a single thread and higher throughput personally. Overall of course still better to wait for next gen devices if you can.
- 4mo ago
- saghm 4mo agoThis is basically my experience as well. I have a moderately recent but high spec desktop (Radeon 6900 XT with 16 GB VRAM, Ryzen 9 7900X 12-core, 64 GB system RAM), and I tried out some recommended models with ollama a month or two ago. Anything not geared specifically towards coding seemed to struggled with actually making tool calls instead of just stating the actions they would take without making them (and trying to get help from them to explain what I needed to configure to change that behavior was useless; qwen refused to believe that it was running in ollama and insisted that it was running from the Alibaba cloud without access to my local system), and the models intended for coding were barely thinking faster than I could type (if they had any ability to show thinking at all). The best "free" experience I've found is using OpenCode with Big Pickle. It's not especially smart, so it often won't produce the correct result the first time, but the free tier is generous enough that I don't think I've hit the limit more than twice over around a month with frequent multi-hour sessions. If running locally is truly the goal, it's not going to fit the bill, but if the goal is just "get the best experience without having to pay for a sub or tokens", it's the least bad option I've found so far.
- alexpotato 4mo ago> The best "free" experience I've found is using OpenCode with Big Pickle. They now offer DeepSeek V4 Flash for free and it def feels like a step up.
- rapind 4mo ago> The best "free" experience I've found is using OpenCode with Big Pickle. I have absolutely zero interest in free. I honestly don't think I'm even remotely in the same demographic as people using free tiers / models. I want to pay. I don't want my data used for training. I want it to be open. I want it to be consistently up (more than Claude!). I want it to be fast. I don't want it to be subsidized as that's just an excuse for shitty quality. Deepseek flash knocks it out of the park on all of these except you're data is used in training. I'm fine with it being hosted since there's no way I'm using it 24/7, but data MUST be private. Basically I want Hetzner and OVH to run open model clouds. I'm convinced this is going to happen eventually when everyone realizes this is a commodity.
- iwontberude 4mo agoThey are good if you were clever enough to buy a powerful enough rig before memory went up. For everyone else I say just wait. M1 Ultra 128GB and higher is sufficient to run gemma4:31b-mlx or qwen3.6:35b-mlx with subagents. It’s only slow if you don’t know how to plan your work effectively.
- aftbit 4mo agoIMO running local models "well" still requires an expensive hardware investment. You really want 96GB of VRAM on a modern Blackwell arch to run these models with decent KV cache. Trying to run them on a unified memory Mac, an AI Max AMD processor, or a DGX Spark-alike is really just asking for trouble. Prefill kills perf. If you throw the right GPUs at the problem, they become much better - but still not quite in the realm of Sonnet or DeepSeek 4 Flash, let alone Opus / DeepSeek Pro or Mythos/Fable/GPT-5.5. Given enough budget, power, and cooling, you can run some pretty good data pipelines, but for code, I think it still makes sense to shell out to an API provider most of the time.
- eek2121 4mo agoNot really, Qwen 27b offloads to a decent gaming GPU (RTX 4090 in my case) without needing tons of RAM.
- mathisfun123 4mo agocan you give more info? llama.cpp vs vllm? config? i wanna try specifically this model
- maldie 4mo agollama.cpp to get 115 tok/s on RTX 4090 with Qwen3.6-27B. For example in Windows the latest CUDA variant llama-b9678-bin-win-cuda-13.3-x64.zip and Unsloth UD-Q4_K_XL MTP gguf: llama-server.exe --host 0.0.0.0 --alias "Qwen3.6-27B-MTP" -m "F:\Qwen3.6-27B-UD-Q4_K_XL-MTP.gguf" -c 75000 -ngl 99 --metrics --temp 0.6 --top-p 0.95 --min-p 0.00 --top-k 20 --presence-penalty 0.0 --no-mmap -t 16 --spec-type draft-mtp --spec-draft-n-max 3 --reasoning on -fa on --parallel 1 -lv 4 Note that this does not use kv cache quants as in my case quants offload to CPU and tanks performance. Also keep in mind this almost maxes VRAM usage so any additional browsers or other programs that use VRAM should be closed. For chat go to http://localhost:8080/ and minimize the window to maximize perf as the web page UI draw itself consumes a lot of GPU perf via constant context switching. Can try bigger than -c 75000 until perf gets lower than 100 tok/s - that means something is off as windows starts paging out memory or other issues. -c 50000 seems sweetspot if running browsers and stuff that consume 2GB VRAM. If wanting more than -c 140000 then likely need to use a bit smaller model quant. CPU usage should be near zero, maybe 1 core load. If you see 8+ core load then settings are off and something is offloaded to CPU (for example kv cache). GPU load should be about 100%, meaning it utilizes work optimally in this case. -t 16 can be omitted or set to the amount of physical cores, not important in this dense model that is 100% in GPU. Can be pushed to 125 tok/s with that model if using --spec-draft-n-max 4 but VRAM usage also increases, so context needs to be smaller. If speed is not important and want max context length then remove the draft-mtp parameters and also might need to use k and v quants like --cache-type-v q8_0, leave k f16 if possible to keep quality.
- everdrive 4mo agoWhat counts as a lot of memory? What could someone do with 16 GB of RAM?
- ValdikSS 4mo agoGemma e2b, Gemma e4b. It's made for smartphones basically. You can run e2b with 8GB RAM.
- monegator 4mo agogemma runs pretty well
- zozbot234 4mo agoModern inference engines can stream in weights from SSD in order to save on RAM, but this makes inference very slow, especially for the trivial single-session case. (Jury is still out on whether batching multiple sessions together can mitigate this well enough, but even then that's mostly helpful for the "running lots of inferences overnight and getting fresh results first thing in the morning" case. Which is interesting (the big third-party suppliers don't really offer a way of doing this at reasonable cost) but a bit of a niche.)
- trouve_search 4mo agogemma 12B 4bit quant; try something with MTP and an AWQ quant
- abalashov 4mo agoNot a ton. I'd say 64 GB minimal to play, 96-128 GB better.
- throwawayffffas 4mo agoNah, you can run the 24b - 35b class with between 90k and 256k of context with about 40GB and they are pretty good. Especially the MOE variants fit neatly in 40GB.
- 4mo ago
- dominotw 4mo agomaybe painful if you are using it like a chatbot. you are sitting there waiting for response. vs ambient ai like automatically classifying your family pics and discarding random things like parking floor number pic. i use it usecases like that latter and they are fine.
- ridiculous_leke 4mo agoA median laptop is no bueno for running a reliable model(which will be qwen 27b as per my reading here and r/localllama). Powerful macs would be prevalent in certain areas of the world but in rest of the world personal machines aren't always that powerful.
- FuriouslyAdrift 4mo agoKimi 2.6 or 2.8 is what we are playing with locally. They need 512GB to 1TB to run with full capabilities so that's not exactly "desktop" Our GPU computer server cost $110k.
- abalashov 4mo agoBut boy, it must be glorious. I use these through OpenRouter and rarely bother with Claude anymore.
- atomicnumber3 4mo agoI largely don't disagree with you but come to a different conclusion. I have two systems: 1) a "programming desktop" with a $500 upper mid range Ryzen (idr exact), 8GB VRAM Radeon card I bought solely for RuneScape, and 64GB ram 2) a maxed out Alienware 16 Area51, so it's a 5090 with 24GB vram and 64GB system ram. I bought it for gaming, of course. I run qwen 3.6 35B A3B Q6 with 200k context window. I compare this to Claude pro max or whatever that I use at work. The main difference between the machines is that the one with the RuneScape gpu does 10 TPS while the Alienware does 30-40tps. Both are fine though the 30-40tps is obviously a lot snappier. I find with both models that: - they do really well at "be a 30GB zip file of reddit and stackoverflow answers" - they do really well at point fixing random bullshit errors that would otherwise waste my time (this is related to above of course) - they do quite well at, given a pretty good specification of what you want, figuring it out, even if you've specified several steps needed - they both cannot really be given a large ish task and left to just drive it on their own The main difference between the two is with that last one, Claude is somewhat better and figuring SOMETHING out, but if Claude is having to figure it out, it's probably because I don't know what I want and it's very likely to not make a sane choice, and will generally produce slop given even the slightest amount of leash still. I've also found that the boundary between "well specified small to medium thing" and "idk just do thing and figure it out" is the difference between you keeping control of the code and losing control. There's an "escape velocity" of AI use that, when you hit it, you're doomed to slop forever. (Or you have to deorbit... enjoy that). And while claude might have slightly higher velocity allowed while remaining suborbital, it's very diminishing returns. So, are these models "worse" than Claude? Yeah. Am I looking forward to continued improvements? Yeah. But I now also have no desire to pay anthropic any amount of money, which has the nice side effect that i won't be helping them end up with so much money that they can distort our democracy.
- hnlmorg 4mo agoTo be honest even the cloud models are a hot mess at times. This week I’ve spent more time rejected code from OpenAI models than I have approving it. In fact it really feels like OpenAI models have taken a nose dive this week compared with Claude. At least for my specific workloads (these things are so variable it’s like trying to compare Google results…)
- citizenpaul 4mo agoThey are still terrible at tool usage which loses 99% of the effectiveness of the agent. I've had to concede and use paid frontier models that can use tools or its not worth using agents....copy...paste....copy....paste....
- iwontberude 4mo agoYour models aren’t big enough and they are forgetting about the tools. Try a larger model. If you can’t, then your rig was too underpowered anyways.
- freehorse 4mo ago> You have MoE models (gemma 26b, qwen 35b, north mini code 30b) who are pretty fast, but make a lot of mistakes This is sadly also my experience. I wish we had some MoE models with a higher ratio of active parameters per total. My experience is that the newer MoE models that can run in a 64b laptop have too few active parameters to be useful outside narrower, specific tasks. Mixtral 8x7b was a 14b active parameter (56b total) MoE model a few years ago and was probably the best model one could run in that range for some time, but it is too old now. I have been using the qwen 27b and it is great, but running a dense model like this in a macbook is a bit suboptimal, and i wish I could run sth faster than 15 tok/s.
- c0rruptbytes 4mo agoI would try a 6-bit MoE and maybe with unsloth's studio, they claim to have auto tool fixing which is where i see a lot of issues with MoEs I'm on a 48gb M5 Pro right now and it's been okay, a lot of my rough experiences have been with MLX and I'm finding that GGUFs are okay now
- Stagnant 4mo agoI've been using unsloth/gemma-4-31B-it-qat-GGUF daily for various small parsing and programming tasks using opencode and llama-server's front end. The past couple of weeks have made a big difference after google released the QAT variant and llama.cpp got support for MTP which means it is possible to now get 60-80 Tok/s with RTX 4090. The model fits in VRAM comfortably enough to keep it loaded even while browsing and having multiple programs.
- amdivia 4mo ago110+ Tok/s as another data point on the RTX 5090 (Gemma 4 31B QAT + MTP at UD-Q4_K_XL) (at peak used 27 GB of vram) The real lovely thing was getting 300+ Tok/s (Gemma 4 26B QAT + MTP at UD-Q4_K_XL) (at peak, I think I saw vram usage reach 21 GB of vram)
- lmedinas 4mo agothe problem of that setup is that it will run out of context pretty quick. So for coding agent it will limit your workflow very fast.
- robomartin 4mo ago> On top of that, your laptop becomes a loud hot churning machine, it's uncomfortable to work with. Laptop? OK, I've made that mistake before. I understand modern laptops are powerful, but nobody wanting to do serious AI/ML work should be using a laptop for anything other than SSH or similar low-performance access into a proper system. Years ago I fried two laptops just doing finite element analysis work running 18+ hours per day. It was one of those "I'm giving you all she's got, Captain!" workloads. They fried, even with powerful fans cooling them. I should have known better. Such workloads belong on purpose built systems.
- deleted 4mo ago[deleted]
- smcleod 4mo agoThose dense models are pretty fast with MTP now. 40-70TK/s depending on your machine, that's faster than cloud models (although not as smart obviously).
- beadw 4mo agoI think you’re spot on. In my experience people confuse a models ability to solve some benchmark as a sign of its usefulness. Token throughput is often just as important from my personal usage. I am excited for more diffusion models to see how progress happens there.
- peterlk 4mo agoYes to diffusion models! Combo pipelines of generative and diffusion models have super interesting potential
- devilsdata 4mo agoJust to piggyback onto this comment; has anyone tried running multiple of these in conjunction? For example, having a Python script that has one of these orchestrate others, and offloads certain tasks to better/more powerful models, or even cloud models?
- pizzafeelsright 4mo agoyes but then that defeats the purpose of 'local' and if remaining local, the hardware required to run multiple poor models could be better spent running better models. I have attempted to orchestrate using different models, loading and unloading, but the speed is not there and by the time mistakes are discovered considering the lack of quick iteration the results become worthless unless the task is trivial.
- andy_ppp 4mo agoI wonder if it is better to have a machine somewhere running a model for you maybe shared with a few others. I could probably justify a M6 Mac Studio with hopefully 256gb RAM and have a few people all with access to one agreed upon model. I think maybe laptops are too warm and clunky for this.
- wgd 4mo agoThe problem is that the moment you introduce shared remote hardware there's a slippery slope leading right back down to "just pay an inference host for model tokens". If you're transmitting your prompts over the internet to a trusted host you might as well just let that host be DeepInfra or together.ai or one of the many other providers already in that business.
- andy_ppp 4mo agoI dunno, I probably need the web to be able to do work so why does it matter - taking the simple case - of running just myself on a Mac Studio at home or cooking my self on the go I'd probably rather have a cheaper laptop and dedicated hardware. I think for many this is about having control over the model and not about farming things out to a SAAS... what does the saying say opinions are like again.
- not_kurt_godel 4mo agoI had some local model FOMO, trialed for a few days, and tentatively arrived at the same conclusion. I can get a better ROI on the time I spent waiting and dealing with poor quality by just programming by hand myself instead.
- DiabloD3 4mo ago[dead]
- locknitpicker 4mo ago> I don't know about good, I use a lot of local models and they're still pretty painful to run locally You are somehow assuming cloud-based models are not painful. I can tell you my past experience. I was using GPT 5.5 and Claude Opus interchangeably and I prompted them to implement a feature. I paid attention to the agent window and it was literally screwing up implementations, causing tests to fail, and going into test-fail-fix loops to clean up after itself. After a few minutes, it finally called it done. That run cost $0.60. I went to review the code and only half of the source files complied with the instruction files. I prompted the model to clarify why it failed to comply with the instruction file. The model outputs "you are right, I should have complied with the instruction files. That prompt cost $0.30. I prompted the model to proceed and apply the instruction file prompts. It went ahead and applied changes. Success. It cost $0.16. I reviewed the code again. Only half of the sloppy code was touched up. I prompted it to fix the whole mess, not just a couple of files. It complied. One coin less in my purse. So, around a third of the cost of a feature is spent on the model cleaning the mess it left in it's wake. And this was a tiny feature with a plan, a solid set of instruction files. Very expensive. Are costs going down? I doubt so. OpenAI seems to still be spending 3 times it's revenue already. In comparison, local models sound very good.
- NamlchakKhandro 4mo agoPi mono is king. Everything else is hypetrash. If I can't customise it then I won't waste my time using it it getting use to it. Claude code is trash, it's customisability is extremely shallow, open code, codex, copilot, Kiro, etc etc... all trash. Yes even open code.. If open code was so awesome then open claw would have been based on it... But it wasn't. That's should tell you everything you need to know.
- EnPissant 4mo agoWhen running on a GPU, dense models are shaping up to be the best way due to two things: - Maximum intelligence per VRAM (you dont have much) - Dense models can benefit from MTP to get an almost 2x speedup in decode (ie, a 27b dense model with mtp decodes at about the same speed as a MoE model with 14b active param model would). This is important because local llm rarely has parallel streams to batch together. When running on large unified memory like Strix Halo or Spark Dgx, MoE models are usually best: - You can get similar intelligence as a smaller dense model with fewer active params (to compensate for the slower memory) by throwing ram at the problem.
- zozbot234 4mo agoThe problem with batching local LLMs is not any inherent lack of multiple parallel sessions, but rather that local dGPUs lack the VRAM capacity to host KV-cache for several of those at once, whereas unified memory platforms broadly lack the compute headroom compared to memory bandwidth that would actually make batching useful. (SSD streaming a larger-than-RAM model "solves" that latter issue very nicely because it radically slashes the equivalent to memory bandwidth so any saving on that becomes highly significant.)
- nullc 4mo ago> This is important because local llm rarely has parallel streams to batch together. I think most people using agent-like usage could easily run any number of parallel streams pretty often, but you run out of vram for multiple KV caches, unfortunately.
- iLoveOncall 4mo ago[dead]
- xlii 4mo ago> I use a lot of local models and they're still pretty painful to run locally. This really depends on how and what you're using. e.g. I can't suffer through slowness of inference on Macbook but I have gaming rig with quite powerful GPU and I squeeze ~130 t/s on Gemma or ~70t/s on Qwen. Tuning is not optional as well. Qwen on temperatures > 0.5 is unusable for coding and I found sweet spot around 0.32 for coding. Speculative decoding on Gemma4 26B is a 30t/s difference between non-speculative. The worst thing with local models is that I can't just give you a recipe, because what's the best params depends on your use case. In the nutshell I'd compare local models to running game rig on Windows vs Linux. Linux works great if not better than Windows gaming, but you need to embrace some tweaking in order to get there. Is it there? It's not SOTA, that's for sure, but it's working reasonably well.
- onel 4mo agoAgree with this. Open models are the future but currently they are a pain to run locally. As painful as it is to admit, the future might be cloud inference from a trusted provider.
- chrsw 4mo agoThe very understandable desire to not have to rely on huge, centralized companies or powers for tokens has clouded people's judgment on how well these local models actually perform. They've improving, which is great, but for real work I use the best models available right now because they're so much better than local models.
- markdog12 4mo ago100% agree. I've spent many hours testing out local models/harnesses. So far, they're very much not worth the tradeoff. Obviously, I hope that changes.
- adam_patarino 4mo agoI always find it amusing when people would rather spend $200 / mo than let their laptop fan turn on.
- segmondy 4mo agoI run 27B at Q8 with fp16 KV cache at 50tk/sec on 2 3090s. Not 4090, Not 5090. 6 years old GPUs.
- naikrovek 4mo ago> You have dense models (qwen 27b, gemma 31b) who are pretty smart, but pretty slow slowness doesn't matter a lot to me, at home. I will type up a prompt and submit it and let it run while I do other things around the house. I have all kinds of things to do, and most of them do not require sitting in front of a computer. of course faster would be better, but it's not always a requirement. smart and slow is far better than dumb and fast or even nothing at all.