7 ms·
Show HN: Sunk Cost – How long until a local LLM rig pays for itself?
I kept hearing "just buy a Mac and run models locally, it pays for itself" and wanted to check. Sunk Cost takes a machine, a model and how many tokens you use a day, and works out how long the hardware takes to pay back against renting the same model by the token.
Obviously there are other reasons to buy your own hardware aside from just saving money on llms but this is just looking at it from a raw cost saving perspective.
If you have any ideas on how I can make this more helpful lmk!
- mcone 17d agoThe idea that you need a new machine is pretty ridiculous. I bought a used HP Omen with a 3090 last month for $2k. 57t/s with Qwen 3.8.
- rlindsey123 17d agoI've not heard of others running HP with it. Hows much RAM do you have?
- mcone 17d agoThis particular machine has 64GB, but the model is on the RTX 3090 with 24gb. Context is 156k with Pi mono.
- usernomdeguerre 17d agoAgreed, I was also annoyed that the only params on the site were mac products. I run qwen 3.8 on a 12 year old asus and a 3090, 50tok/s. It's not even the only guest running on the box. For my usage profile (not running it 24/7) it's actually less expensive per-month than claude subscriptions.
- redox99 17d agoI'm so happy for the two used 3090s I bought for $500 each after Ethereum mining ended. I even saw them for like $430 at some point lol.
- hyperhello 17d agoI doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
- gruez 17d ago"If they are selling it for less than it cost to make, buy as much as you can." -- Warren Buffett
- taraindara 17d agoOnly caveat is you’re buying time. Not a physical good. It’s only worth what you’re able to get out of it in that time.
- rlindsey123 17d agoIs that a real quote? Golden if true
- tyre 17d agoFor their current models, served directly from their infrastructure, they are profitable after training (which all present models are.) I don't know when we'll have an open equivalent to Fable, let alone whatever (insane) hardware you'd need to run it locally.
- deleted 17d ago[deleted]
- epistasis 17d agoMore than that, running hundreds of conversation streams at once is essentially the same cost as running a single conversation. And then you add on the secondary benefit of having the GPUs running nearly all the time rather than mostly idle... Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.
- 16d ago
- jrflo 17d ago43 years to break even on Qwen 3.8 at 25% the speed of the API, lol. I like the idea of local models for really small tasks like automation/toolcalling, but it will probably never make sense for coding. I tried them and it was just excruciating compared to what you get for $100 a month from a subscription.
- rlindsey123 17d agoYeah it's surprising how long it would take to get back on those local models!
- itake 17d agoI have a home server running vibed applications. VPS host would cost $25/mo or $300/yr. Mac mini can also build iOS applications. I think if you’re a mobile dev, you can have concurrent builds for your agents instead of everyone waiting on a single machine to finish.
- rlindsey123 17d agoWhat models are you running on it? I'm also an iOS dev but I find I need more frontier models to get good quality code from it.
- itake 17d agoI only use frontier models to vibe code iOS apps, as I'm not an iOS developer. I haven't tried the local models post qwen coder 3.5 release for all the reasons. AFAIK, a limiter for iOS engineers (and AI agents) for concurrent feature development is the xcode environment and hardware limits. BE engineers can easily have 3 agents working on 3 different microservices (or gitwork trees), but iOS devs can basically only manage one version of the code at a time, due to externalized state (like derived data and bundle ids).
- ProjectArcturis 17d agoLocal LLMs are not really about saving money, they're about autonomy. Choose the exact model you want, fine-tune it if you want, and no one can take it away from you.
- rlindsey123 17d agoTrue - definitely agree!
- poincareball 17d ago[dead]
- txrx0000 17d agoIt pays off instantly, because OpenAI/Anthropic can no longer see what I'm doing and that's worth a lot of money to me. If I am offloading some of my thought processes to a machine, I want to own that machine. And if I finetune the model, I can gain access to parts of thought space that are cordoned off by OpenAI/Anthropic/Alibaba/whomever due to their "alignment" efforts (i.e. alignment to the AI company rather than me). Otherwise, it's like if someone else owns a part of my mind and has a backdoor into my mind.
- no-name-here 17d agoYou can't run recent openAI/Anthropic models locally anyway, so wouldn't a better comparison be a different provider running Qwen or similar model? As then you can also compare against the exact model you'd have locally and any different data privacy of that particular provider?
- txrx0000 17d agoTechnically true, but the delay between local and closed frontier is only a few months. And individual sovereignty / digital bodily integrity is almost priceless.
- selectodude 17d agoLocal frontier costs a half million dollars to run locally in anything higher than basically ternary.
- txrx0000 17d agoOkay, that's technically true again, but local mid-tier like Qwen3.8-27B is only a year behind the closed frontier. I'm personally willing to be behind by a year if it gives me mental sovereignty against the big AI companies. They are extremely misaligned with me.
- koito17 17d agoGP's point is about "sending tokens to someone else's computer" versus "keeping the tokens locally". I think model capabilities are secondary. In May of this year, I was running qwen3.6:35b-a3b on my MacBook (bought in 2024). Obviously not as fast as, say, running a model on Cerebras, but a year ago it wasn't really feasible to have a local model running on my 2024 laptop with vision support. (Concretely, I was passing apartment diagram pictures to Qwen and making it compare different apartments for which ones would feel the most spacious while optimizing for initial moving costs and other factors.) This was back in May and I wouldn't be surprised if there have been significant improvements since then. Overall, I think it's fair to compare a workflow like "use llama.cpp locally to upload some pictures and ask questions" to "open the ChatGPT app, upload pictures from your phone, and ask questions". Sure, you can't run a model like GPT-5.4 locally, but the model is mostly an implementation detail here. What a user will care about is: "when I go with the llama.cpp option, am I getting useful information from my conversations?"
- shadowpho 17d agoI like this calculator but it’s really wrong at least for dgx spark. I have one and I get 4x the tokens/s .
- rlindsey123 17d agoAw very interesting! This is great feedback - what model are you running? I'm keen to do more crowdsourced data as time goes on.
- shadowpho 16d agoThe big three :) Qwen3.8-flash-next Deepseek4-0731-flash Glm5.3 The latest unsloth llama.cpp has a lot of nice features that runs them faster than before. I’ll have to double check which one runs how fast, but it’s generally 20-40 t/s. (And infil is fast but not sure how that’s counted)
- chasd00 17d agoNot a fair comparison really. If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. That has value a subscription does not. Idk about the quality of this setup but just pasting it here as an example. https://explainx.ai/blog/heretic-llm-abliteration-guide-2026 https://explainx.ai/blog/heretic-llm-abliteration-guide-2026
- ChickeNES 17d ago> If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. When does the average person actually need to do that?
- jerf 17d agoWe've already seen frontier models refuse to answer almost any question that touches on computer security and be very likely to kick out biology and chemistry questions even if they aren't all that close to breeding dangerous viruses or making explosives. I expect this is only going to get worse. "Censorship" isn't just going to be about who you vote for and which political party the model will say nice things about and which it is more likely to say bad things about. It's going to become about whether the hoi polloi are allowed to have effective AIs at all. Like the 1990s internet, AI has outrun a lot of power structures but that is not going to continue indefinitely.
- ChickeNES 17d agoSo you want to remove valid safeguards? And stop misusing the word censorship.
- jerf 17d agoYou asked a question. I gave you an answer. I seriously doubt that if you and I sat down together at a table and banged on this for an hour that we would come to the same definition of "valid". Ask 10 people, get 12 answers to that question. There's going to be a lot of motte & bailey in the next couple of years, where I just want an AI to answer questions about whether my code is vulnerable and people like you will be "Oh so you want an AI that can hack the Pentagon do you?" and it doesn't look like we're going to be seeing eye to eye on that one.
- ThunderSizzle 17d agoClaude Code is $100+ or else be constantly throttled. My usage on GHCP was gonna be $300+ a month. I paid $1350 and threw an R9700 in an existing machine. That's a 4 month pay off or so. Plus, I can feed it sensitive data all day and not be worried where it's going.
- no-name-here 17d agoAn R9700 has 32 GB RAM. Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?
- dwb 17d agoYou should be comparing the value you get. If you get as much value from a local model as a hosted one, the size difference doesn’t matter.
- no-name-here 17d ago>>> Claude Code is $100+ or else be constantly throttled >> Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally? > You should be comparing the value you get But the GP commenter specifically compared the cost of solutions such as Claude Code against a 32 GB model. If they are going to compare cost, they should compare to the cost of a hosted ~32 GB model. Or if privacy trumps everything for them, then just say that and don't bother comparing costs of incredibly disparate solutions, as Claude Code costing $100+ a month was a red herring if they're happy with 32 GB model output - they could have compared to a far cheaper option that matched their local model's quality. It would be like someone saying they were able to buy a bike to get to work, saving them $x million compared to buying a Bugatti. When really, if they're going to compare cost they should compare to a cheap car, or not bring up the cost of an expensive car at all if exercise trumps everything else for them.
- ThunderSizzle 16d agoWell, I don't see a value issue of using Qwen3.6 27B vs Sonnet 4.6 (not sure about 5 yet) I still have to use GHCP at work, and I self-host at home, and aside from the fact self-hosting also forces you to tinker, optimize, etc. - there's not a huge difference in my end result in end user results. I spent quite a bit of time trying to optimize llamacpp and compare 35b to 27b, etc. I don't compare models that much at work. I guess the other part of it is I didn't really know much about cheaper cloud models, but I was attracted to the idea of no longer renting against Claude code, etc. I figured if I could run something functionaly similar from my bedroom on a normal outlet, then all this talk about data centers needing to be built everywhere in the news cycle is obviously just plain stupidity and hype. It appears I'm using about 20.4/7.6 million in/out tokens a month, or on open router, about $20/month. That puts $1350 at a 5-6 year break even (thanks to cheap electricity), I guess. Beyond that, running on localhost as a nice feature of 0 no latency when doing rapid tool calling
- bix6 17d agoFun feature: can you show some sort of list of the best combos? Eg shortest payoff time for best capability in various situations.
- rlindsey123 17d agoGood idea, pretty crude but it's up: https://sunkcost.ai/best/ https://sunkcost.ai/best/ For each usage level, it lists the quickest pay-back in each capability class, with each model on its quickest machine and one click into the calculator to change the assumptions. Short version: at 1M tokens/day the best Sonnet-class option is Qwen3.8 27B on a Mac mini M6, 8.3 years. It only drops under a year if you're running agents at around 20M tokens/day.
- bix6 17d agoThat was fast! At 7 tokens/s (Mac mini) you max at 600k/day so you couldn’t hit those higher amounts like 4M where it says 2 year payback?
- QwenGlazer9000 17d agoYeah no it does not pay for itself just comparing to cloud. Not at these prices at least, people far richer than you or I buy these things wholesale, no scalper, bought a significant amount at cheaper prices, and are wired up the ass with VC money. The premium is not having your million dollar prize and career stolen by billionaires.
- rlindsey123 17d agohaha 100%. We used to just rent our homes. Now we have to rent our intelligence
- deleted 17d ago[deleted]
- deleted 17d ago[deleted]
- jocelyner 17d ago[dead]
- monksy 17d agoI wish you could put different setups on here. I have a couple of A6000s on an AM5.
- rlindsey123 17d agoI'm keen to add a way for people to add community based reporting which would allow this. Would you want to see anything else on the dropdowns to be able to enter your data on?
- bpbp-mango 17d agocan you add RTX cards too please? 5090 and 6000
- gfody 17d agoshould throw in a tt-quietbox
- harhargange 17d agoAlso, i also use my gpu for rendering and learning and playing games.
- serial_dev 17d agoIn the “The small print that isn't small” you describe all the disadvantages of running your models locally, but none of the advantages (just check the rest of the comment section for inspiration on that).
- 01HNNWZ0MV43FF 17d agoI was just gonna throw a beefy Ryzen into an ATX chassis. I don't want to pay Mac prices
- Zetaphor 17d agoThis tells me that the max throughput for the models I'm running on my hardware is lower than it actually is. Please allow us to tweak all the variables instead of locking me in to whatever rate you found by searching
- rlindsey123 17d agoThanks for all the feedback. You can now enter your own measured tok/s for any machine and model.
- v3ss0n 17d agoBesides from privacy: I already making twice now.you own the hardware and the price had doubled since i bought. Almost tripled. You missed the opportunity and i have 4 of those awesome machines. Cry on. I sell those to business who need local air gapped requirments and I make a lot more money! I can run the alliterated models where none of the service prvoider even dare to provide. THose benefits outweights a few K. And show me an api provider that allows me to run 10x agents concurrently for 5 days straights .
- ChickeNES 17d ago> And show me an api provider that allows me to run 10x agents concurrently for 5 days straights . Any of them on a Max/Pro plan as long as you are smart about model selection? That's my main objection to local inference, I'd need a whole rack of GPUs to do as many things in parallel that I can do for $400 a month. I do plan on setting up some local inference hardware, but...RAM and GPU prices alone are $$$$
- 0xbadcafebee 17d agoIt pays for itself very quickly if you do 24/7 generation. Use an AI agent that orchestrates other agents working on many things at once constantly. If speed is a factor, you'd not buy a Macbook, you'd buy dual RTX 3090s. About the same price, but at least 6x faster than M5 Max. The benefit of constant generation is you can do a lot more research, coding sub-agents, experiments, etc in parallel when you're not "at work". You end up getting a lot more work done than if you only sit there babysitting sessions.
- jayle 17d ago[dead]
- dumberquestions 17d agoYeah I'm surprised no one pointed this out, if something like persistent agents gets more popular/useful, the local option pays for itself surprisingly quickly.
- redox99 17d agoThe math is wrong, the tok/s is at least 2x that, at least with MTP and Q8 KV which you should always use. And the default tokens a day is ridiculously low at least for coding. Having said that, it will never pay for itself. A simpler more absolute math is, if I buy a Mac and use it to sell tokens on OpenRouter, will I make a profit? And the answer is no.
- afarviral 17d agoI want the autonomy but local models of the size I would have the means to host wouldn't be capable enough. What usecases tend to suit these smaller models that tend to produce incorrect or otherwise flawed responses often? Could they work for anomaly detection and what would a rough architecture look like?
- rubyn00bie 17d agoThis is a bit weird because it automatically changes the model depending on the amount of VRAM available, and some of the smaller models are more expensive (presumably because they're being provided via OpenRouter by someone with some GPUs in a colo or smaller providers). It also doesn't allow changing the tokens per second (my 5090 can get like 75-130 tokens per second [assuming I can fit the model in RAM]); which, then results in woefully under-estimated limit on how many tokens a day I can consume. Some improvements that I think would make this more useful: 1. Allow manually setting tokens per second, or as an alternative, let me jack up the number of tokens a day. 2. A sort of backwards flow "if you want to run this, at X tokens per second, with Y context, you'd have to spend Z." 3. Add support for configuring multiple RTX 6000 variants. When I was making heavy use of DeepSeekV4-pro I was burning somewhere around 1.5 billion tokens a month, and that was just using it in my free time on random projects. It was something like $24 at the time because of the initial discount/promo period. I don't think there's anyway in hell I could ever run that (on current hardware) for less money. I think the calculator shows from a purely financial standpoint what we all know... that yeah, it's definitely not worth it if money is your only concern. That (cost per token) will eventually change. Models will get better, more efficient, VRAM prices will come down, VRAM capacity will rocket upwards, and the economics of it all will change. It would just be really cool to have the calculator show me exactly how cheap they'd have to get for it to make sense. I need to finish up some work and make dinner, and if no one else beats me to it (anyone is welcome to) I'll ask Fable or Opus to knock that out.
- kjshsh123 17d agoOn a purely monetary basis it probably never will. You're competing against companies that get tax breaks, locate themselves optimally, and have large economies of scale. Also, if it did, the hardware would be bought up, raising the price until there was no economic profit again. If you can find a unique application for it then maybe?
- binary132 17d agoUnstated key concept: “At today’s prices”
- binary132 17d agoThe real question should be why anyone would voluntarily continue to spend money on a software service that costs as much as an expensive computer when they could just buy an (upgradeable) expensive computer and use it as much as they want, approximately forever. Imagine owning nothing and being happy.
- TokenLat 16d ago[flagged]
- nickyocean 16d ago[flagged]
- itsmeduncan 15d ago[flagged]
- kedarpotnis 14d ago[flagged]
- Sky_Joy3 14d ago[flagged]