5 ms·
Yes. Llama.cpp + Qwen3.6-35b (MTP) + OpenCode is quite capable and runs on a single RTX 3090 and is faster than most cloud models. Quality is like running edge
by pierotofy 4mo ago
Yes. Llama.cpp + Qwen3.6-35b (MTP) + OpenCode is quite capable and runs on a single RTX 3090 and is faster than most cloud models. Quality is like running edge models from 8-12 months ago. Setup details at https://github.com/pierotofy/LocalCodingLLM/ https://github.com/pierotofy/LocalCodingLLM/
- atomicnumber3 4mo agoSame. I have no desire to use Claude at all anymore.
- pierotofy 4mo agoYep. Screw Anthropic, CloseAI and all other rent seekers in this space.
- akulbe 4mo agoI have an M2 Max MBP with 96GB of RAM. What models and setup would you use for this kind of configuration?
- monirmamoun 4mo agodownload LM Studio to play with, and it will let you search for models... try Qwen3.6-35B-A3B at 4,5 or 6 bits (6 bit XL is near perfect) and use pi coder or another harness to access it... you can also try Unsloth studio and try same model to start. LM Studio slighter easier to use, Unsloth probably better quality. Neither one is super great quality by the way (meaning: they crash or act weirdly too often to be full production solutions, but can work for local coding). ONCE YOU DOWNLOAD EITHER APP... it will let you search huggingface for the models. Just type qwen to start looking and ... start messing around. And you connect the pi coder harness using the http interface that LM Studio and Unsloth offer to the engine API, so make sure you figure out that url and turn it on... something like 127.0.0.1:1234/api would be a typical IP (localhost) and port (1234 is used by LM Studio)
- jacobgold 4mo ago"Quality is like running edge models from 8-12 months ago." That sounds great for hobbyists but IMHO it wasn't until Opus 4.6 was released six months go (Dec 25, 2025) that we had a model good enough for professionals to use as a primary driver of their coding agents. That seems to be the threshold worth aiming for.
- pierotofy 4mo agoI use it for work.
- jacobgold 4mo agoThat's cool if you prefer it, but it is hard to imagine it being a strictly rational choice when much better quality is available at a price that is small relative to the cost of an employee. Or is there something specific about your use-case?
- vector_spaces 4mo agoNot all work requires every facet to be so sharply optimized, and there may be other constraints that are completely invisible to you. Some that were easy for me to imagine: the parent works in a heavily regulated industry, their IT team is slow-moving and paranoid and this is a safe, under-the-radar workaround, the output is "good enough" for their purposes and they find tinkering with it to be fun. Regardless I don't think it's fruitful to be so condescending with such little insight into this person's situation. Even if you had total insight -- let people be and withhold your judgement, or at least keep it to yourself. Making people feel stupid is a great way to turn people off to pretty much anything else you have to say
- lokar 4mo agoWon’t it depend on what you use it for? A less capable system might be fine for boilerplate, moderate re-factoring, etc. Not everyone is building whole features in one go.
- pierotofy 4mo ago
- dominotw 4mo agohow much does the setup cost if i want to buy all the hardware now and increased power costs?
- lelandbatey 4mo agoI use it, it's good, I get work done, but know that they really mean it when they say > "Quality is like running edge models from 8-12 months ago" Don't expect Opus, expect more like Haiku. If you micromanage it, you'll get great results. If you want it to be a human in a box, it'll flounder.
- dheera 4mo agoAm I doing something wrong or has ollama become shittified? I'm looking at https://ollama.com/search https://ollama.com/search and the top few models like kimi-k2.7-code say "cloud" and I can't seem to ollama pull them. I thought the whole POINT of ollama was not-cloud?
- satvikpendem 4mo agoOllama is not recommended to be used. Use llama.cpp.
- toyg 4mo agoYes, you've nailed it. Ollama are desperately trying to pull a Cursor - like 3791 other projects in this space.
- hoherd 4mo agoI experienced the same situation a month or two ago. One of my friends sent me this article that was illuminating. https://sleepingrobots.com/dreams/stop-using-ollama/ https://sleepingrobots.com/dreams/stop-using-ollama/
- jmorgan 4mo agoThe larger models are available on Ollama's cloud as most folks don't have the hardware to run 500B-1T parameter models.
- jubilanti 4mo ago> I thought the whole POINT of ollama was not-cloud? It was at first, then the developers realized they had a massive userbase they could monetize. A tale as old as open source...
- trueno 4mo agoi have a 128gb m4 max macbook pro i've been wanting to tinker with this stuff but genuinely never find the time. any mac users in here running similar to the above that can share their experience? i always see great debates with local stuff but the space is constantly moving goalposts and all the vernacular is pretty unfamiliar to me. i'd love to understand what people with objective experience feel they've traded away (or gained) when going local so i can determine for myself if these things are a good fit.
- htrp 4mo agoUse your ClaudeCode sub and tell it to set it up for you
- brycesub 4mo agoIf you have a 128GB Mac you really ought to try out: https://github.com/antirez/ds4 https://github.com/antirez/ds4 by the creator of redis. This is probably as close to it gets to state-of-the-art local LLM + agentic coding.
- lostlogin 4mo agoThank you.
- trueno 4mo agowell this is supremely interesting thanks for putting it on my radar
- __mharrison__ 4mo agoUsing this just this morning on my DGX Spark. A little slower than frontier models but my $200/mo weekly usage exhausted with 3 days left on the week... (Shouldn't have done that refactoring job in high mode)
- dirkolbrich 4mo agoI have the same machine. You might look into https://omlx.ai/ https://omlx.ai/ a „macOS-native MLX server“. pi.dev for the agent with MCP, web-search and sub-agents extension.
- daveidol 4mo agoDo you do your dev work on the windows machine (referenced in the docs), or do you remotely access it from a separate machine? I ask because I have a RTX 3090 kicking around in a gaming desktop, but I don't use it for any dev work (I use a Macbook Pro).
- snake_n_my_boot 4mo agoI have a similar set up and have been using it to learn and tinker with open models. I run Ollama on the gaming desktop and point OpenCode to it from my MacBook. Works nicely for me so far.
- NamlchakKhandro 4mo agodon't waste your time with windows.