5 ms·
It’s not. Do it as a hobby or for privacy but for performance just use a frontier model api. You’re paying less than cost for something that would take tens of
by pcarolan 1mo ago
It’s not. Do it as a hobby or for privacy but for performance just use a frontier model api. You’re paying less than cost for something that would take tens of thousands to set up locally.
- Gigachad 1mo agoIt does make me wonder how the hosted stuff is so cheap. For pretty much everything else, hosted/rented is more expensive but offers better convenience and flexibility. But for AI, even if you consider the total lifetime cost and are utilizing it heavily. You never break even by buying.
- api 1mo agoThere are economies of scale but there’s also a data center bubble (probably) so there might be some selling dollars for fifty cents going on.
- mrngld 1mo agoHere's the thing that's a little different about data centers; we can tell from Anthropic and OpenAI that they're capacity constrained. Inference demand is there. I notice Cerebras doesn't offer much directly any more, all their capacity is getting completely sucked up by B2B sales. Grok did overbuild, but Anthropic was so desperate for more compute they ate their pride and leased the excess capacity. That means all these data centers are being heavily utilized by actual end user inference demand. Well, some is research on new models, but a lot is actual end user demand. No one has given an explanation of why peoples usage would decline. On top of that, margin on inference appears to be decent. It's model training that's a serious financial burden. And maybe that's where there will be a slowdown, maybe the market doesn't justify spending as much on R&D as it does, but the end demand for inference is there. Does that justify these stock prices? That's a different question. But the housing boom left behind endless rows of empty homes because demand disappeared. The 'dot com' boom left behind thousands of miles of dark fiber that'd been built out well ahead of demand for bandwidth. I can see the stock market having a giant sell off, but I don't see data centers sitting idle in that same fashion.
- api 29d agoThe number of planned data centers and the scale is pretty nuts. Here's what you have to believe: - AI demand is at least several times larger than what can currently be satisfied, or will grow. (This one I can buy, but...) - AI chips (GPUs, TPUs, compute-in-memory, whatever else is being studied) will not get significantly more efficient than they are now. It will not be possible in, say, 5-10 years, to do 2X or 4X or 10X more AI requests per rack than is possible now. I think this one's the single most likely thing to be false, since all computing history contradicts it. - Edge devices (PCs, laptops, specialized but smaller scale AI compute nodes) will never be powerful enough to run frontier models at a reasonable price that's appealing for professionals, enthusiasts, or businesses, and there will never be a market for this. None of the demand will be served on-device or near-edge. AI must all go in giant data centers. - AI models will not become significantly more efficient than they are now. There are no large gains on the table from better model architectures, better training, more efficient quantizations, better harnesses, etc. If all those things are true, than the current planned like 4X-10X increase in data center capacity makes sense. If even one or two of them are not true, then the planned data center build-outs start looking excessive. If all four are not true, it's a total bubble that will crash and burn. Answer is probably somewhere between, but how far toward bubble? That's why I picked a number like "only 20% ever gets built." It might be as high as 50%. It ain't gonna be 100%. The planned built-out is batty. Oh I forgot two more... - Data center capacity currently serving non-AI work loads does not shrink through either reduced demand, more efficient software, or (most likely) faster chips and denser RAM. If that happens, more pre-existing DC space can serve AI work loads. - Orbital solar powered compute nodes never happen. If this happens (free power! much less political opposition!) then terrestrial data centers have significant competition.
- asteroidburger 1mo agoIt's a time sharing agreement, just like old-school mainframes and such. You're not getting a full machine to yourself, but a few cycles at a time.
- srcreigh 1mo agoThey're not cheap at all. I did one xhigh Qwen 3.8 27B agentic coding task last week via OpenRouter and it cost me like $10. 99% of the cost was in input tokens, I only used like 100k ish output tokens. It was a one shot task asking the agent to implement proxy injection to Guice. It did a pretty amazing job. If you were to use hosted LLMs for a lot of agentic coding, a maxed out M5 Ultra Mac Studio would pay for itself in under a year.
- Gigachad 1mo agoQwen is weirdly expensive. Deepseek v4 flash is dirt cheap. You'd need at least 128gb of ram to run this model and in my experience, a days work with it costs around 80 cents.
- srcreigh 1mo agoSo I ran the math, assuming the agent takes 75 turns per 200k context, with deepseek v4 flash it costs around $2.57 to reach 1M context in 375 turns. Cached input costs scale quadratically with # of agent turns. Considering that I hit the 1M compaction multiple times per day with codex, it would definitely cost at least $5-8/day to use deepseek how I normally use codex.
- anotherCodder 1mo agoI've been hosting Qwen3.8-27B myself. On my endpoint it's $0.30/1M in, $0.10 cache, $2.03 out - so those agent turns that re-send the same prefix get a lot cheaper when cache hits. UI at inference.tiyuvta.ai/app if you want to try it. Hosted is up to 210 tok/s and 280ms TTFT with reasoning off.
- whatsThisBtn4 1mo agoI watched someone at a fortune 20 company get embarrassed for buying a Mac to run a 70B model in 2025. He was a lead engineer, so after he announced it wasn't going to work, everyone pretended it never happened. But we all knew.
- copper-float 1mo agoSounds like a really rude workplace. Who cares if he wants to try running things locally?
- whatsThisBtn4 1mo agoWhat was rude? No one bothered him. Also the entire purpose of them buying it was so the department had a LLM. I proposed A6000. That ended up working.
- deleted 1mo ago[deleted]
- ux266478 1mo agoThat's not even remotely close to being true, even once you account for capex. You have to look at the actual usage, look at the token limits. Even if you're paying Anthropic $200k/month for scale-tier, you're going to blow through your token limits trying to run max output 24/7. Three users running Opus 4.8 at max non-stop will probably clean your monthly allowance from daddy Dario in less than a week. With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. It does mean 8 multi-trillion parameter models unquantized running 24/7 without pause. And you get the full month like that, your monthly token limit is the time in a month. That cluster, the electrical upgrade, the cooling setup, and the electricity to run it all costs less in 2 months than your maximum affordance from Anthropic does in the same time period. Two billing cycles, and realistically it's more like two weeks. In 4 quarters you've wasted over a million. Like, what are we talking about here? Now if you aren't using AI all that much, which is perfectly valid, and especially if you aren't using it at its absolute maximum, the story changes. Because even though at that point you're not paying nearly as much in electricity to run the cluster anymore, you still have the $300k+ capex to get the setup in the first place. But if we're not redlining it non-stop, then we're not really talking about performance anymore, are we? If your org never comes close to hitting token limits, it's probably because AI is rather marginal for you. Which again, is perfectly valid. I don't even use AI professionally. Fact of the matter is, if your corp can justify the capex for a cluster and makes heavy use of AI, you are literally burning money by not having one in your building. The numbers are painfully obvious. Even deepseek isn't as cheap. This is before we get into things like LoRAs, custom inference pipelines, etc. which you know are kind of important if you actually care about model performance.
- Aurornis 1mo ago> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized running 24/7 without pause You can’t even run one unquantized multi-trillion parameter (>=2T) model on 8 x MI355x with enough context for concurrent users. I don’t know how you think it’s going to run 8 of them at the same time. Did you mean 8 concurrent sessions? Your math is way off across this post. If replacing an Anthropic subscription for a whole company was as easy as buying a box for the office and then breaking even in 2 months, it wouldn’t be some little secret that we only discover in a comment online.