5 ms·
If I said, “I want a setup that is usable for an agentic coding workflow, and it MUST be local”, what’s the smallest/cheapest option right now? It’s _technical
by equalsione 2mo ago
If I said, “I want a setup that is usable for an agentic coding workflow, and it MUST be local”, what’s the smallest/cheapest option right now?
It’s _technically_ possible to get agents running on all kinds of setups but there seems to be an (undefined) floor for useful setups.
A lot of the stories people have about getting setups running on relatively low end hardware turn out to have huge compromises or run into issues on anything but trivial cases. I’ve found it hard to find a consensus.
Or maybe I just don’t like the multiple thousand dollar price tags people are suggesting…
- timmmmmmay 2mo agostraightforward solution is probably Qwen3.6-27B-Q4 running on a used RTX 3090. price on those is unfortunately high, in fact so high that getting a new Radeon AI PRO R9700 might be a better deal a solid step up from there is anything that can run Deepseek V4 Flash but the hardware ask there is a bit higher
- kamranjon 2mo agoI actually have been exploring this very thing! I think the best option right now, since Apple has raised prices and Mac minis are basically impossible to get your hands on, is to build your own micro-itx machine. I actually built a mini-itx machine, but it does restrict your options a bit. The Arc series Intel GPUs are what I think make this possible. I built a machine with an Arc b50 - it runs Gemma 26b a4b qat at around 30tok/s with their MTP head and prompt processing sits at around 500 tok/s. The really beautiful thing about this setup is the entire energy envelope of this machine sits at 120w at full load - when idle, it's at 40w and i've done some work in ubuntu to basically intelligently hibernate, which drops it to 0 watts when not in use. You can use a raspberry pi and Wake on Lan to wake the machine up for a overall draw of around 5 watts when not in use. All in all this machine cost me 1.4k to build - but if you used micro-itx instead of mini-itx parts you could do it for under 1k - it has just 16gb of ddr5 but you don't really need more if you use models that can fit in vram. I think it's pretty incredible that you can run an actually useful coding agent on a machine with a power envelope that is less than an incandescent light bulb. If you go up to micro-itx you can do even large cards like an intel b60 with 24gb or a b70 with 32gb and run even more powerful models. For all of these intel GPU's you'll want to compile the latest llama.cpp version with SYCL support - they are getting speedups every day, so worth staying on the edge.
- jononor 2mo agoIt depends on your tolerance level for having less than frontier LLM capabilities. First level worth trying, Qwen 3.6 35B A3B with a 16 GB VRAM (example 1x 5060ti 16gb, 600 USD for the card) with partial GPU offloading. Next level would be Qwen 3.6 27B / new Muse Spark / Gemma 31B with 32 GB VRAM (2x 5060ti or 1x 9700 Pro). Third level would be DeepSeek V4 Flash with 192 GB VRAM (2x Strix Halo at some 8000 USD total). These models can be tried on OpenRuouter etc, or you can deploy vLLM on rented GPUs to get a feel for what level you would want before committing to buying hardware.