4 ms·
This should be a warning to those who feel that it's ok to offload your creativity to a subscription service. Always need a local model in some form.
by wasabinator 6mo ago
This should be a warning to those who feel that it's ok to offload your creativity to a subscription service. Always need a local model in some form.
- jazz9k 6mo agoThe 'local model' is called your brain.
- mingus88 6mo agoI’m sorry but that’s just dumb. An LLM is a tool. Your brain is not a substitute for an LLM in the same way your fingers are not a substitute for a wrench. The year is 2026 and if you are using your brain on chore work like one-off scripts, refactoring, boilerplate test code, then you are wasting time and money and I don’t want to work with you. Local models are fine for this and can do it in a fraction of the time your brain will take to even get bootstrapped
- adithyassekhar 6mo agoThe year is 2026 the average RAM for the most common type of developer’s (web) machine is 16GB. 8 will be the lower end. Tell me which model can one run on this machine locally?
- porridgeraisin 6mo agoYou can use free models on opencode. Minimax whatever is just free. It's more than enough for these tasks.
- avaer 6mo agoYou could judge the costs of the AI products you're using by the standard API pricing, not promotional subscription offers.
- glimshe 6mo agoNot even that way, given that the price is still highly subsidized by investors and circular deals.
- kadoban 6mo agoFor me, it's not even cost necessarily. If they decide to change the product they offer, the old one is gone. I refuse to use anything for personal use that's not at least _available_ as model weights.
- brewtide 5mo agoBingo. This same scenario with IOT hardware/software requirements and the ever changing software updates where features get added/removed (and firmware), etc would have so many here up in arms! Oddly, less (vocal) arms on this specific case.
- locusofself 6mo agoAre there local models that are anywhere near as good at coding as opus 4.6?
- jasonjmcghee 6mo agoPeople will insist otherwise, but I haven't seen anything close to sonnet 4.6 that can be run locally.
- Incipient 6mo agoI don't think anyone can honestly say a huge frontier model is actually going to be matched by something running on 64gb locally?
- jasonjmcghee 6mo agoI have read many comments saying Qwen3.5 various ~30B models, Gemma 4 ~30B models and now Qwen3.6 "better than sonnet". I don't know how large sonnet and opus are but the rumor is 1T and 5T respectively.
- urig 6mo agoYou don't have to use the most recent bleeding edge model to succeed. A local FOSS coding agent coupled with a reasonably priced LLM could yield the optimal ROI.
- kadoban 6mo agoNot really. Qwen 3.5 and Gemma and a couple of others are quite good though, and the quants are _very_ runnable on a good gpu.
- rvz 6mo agoI keep telling them and they still want to spend money on tokens at the Anthropic casino, even though they are egregiously price gouging and applying upper limits so you spend more on tokens. Sometimes you can't help gamblers who want to gamble on tokens to hit the jackpot on fixing a typical issue which can be done by local models or even reading the documentation.
- para_parolu 6mo agoThere is very little vendor lock. We can keep using subsidized model until it’s not. Then switch to next subsidized model.
- ares623 6mo agoIt's like chairs!
- ratg13 6mo agoThis doesn’t affect existing users. This is a simple supply and demand curve. Higher demand means the price goes up .. this has been true of things since before SaaS and before computers
- HumanOstrich 6mo agoThanks for all the logical fallacies in one comment.
- avgDev 6mo agoNo. It means eventually existing users will be affected too. These companies are deeply in the red.
- rurban 6mo agoLocal models are not comparable to the FOTA models at all. I know what I'm saying because I do have 4 local H100's in my server, and could run the very best local models. It's night and day. They are unusable and stupid.
- poisonborz 6mo agoFor what do you use the 4 local H100s then?
- rurban 6mo agoFor training our AI model of course. Inference is for the cheaper machines.
- wasabinator 6mo agoNot all tasks require a frontier model
- bravetraveler 6mo agoI get perfectly acceptable results from a Strix Halo PC the size of a shoebox, man. An APU that uses ~150w, has 0 discrete GPUs, and a bill of $0/m. What's more, it doesn't go down every week, limit use, or change the terms at a whim. I'll burn/discard 'frontier' tokens (at work) only because they're mandated and they foot the bill. I'd rather resell them; meet the asinine requirement from $EMPLOYER, provide cover for outsourcing to my equipment, and get a return for the hassle. TLDR: perhaps you're holding it wrong or haven't tried the latest, as we so often hear. That's a lot of GPU for not much utility.
- rurban 6mo agoWell, my python and typescript folks are also happy with the simplier local models. But I'm using more advanced stuff, C/C++ embedded real-time, vision AI, and compilers.
- bravetraveler 6mo ago
- wookmaster 6mo agoI’ve been trying to bring this up at my work, you’re putting all your intelligence into a service you don’t own. What do you do when it’s down or they quadruple the price ?