5 ms·
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage wher
by preommr 2mo ago
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,
I don't have the words.
I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
- buckle8017 2mo agoThey over purchased hardware. This is very likely priced below recovering the cost of the hardware but still above operating expenses.
- infecto 2mo agoWhat evidence is there? I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.
- qntmfred 2mo agosama literally just said they wish they had bought more. the price drops are almost certainly due to good old fashioned hardware innovation (wafer scale with cerebras) and optimizing hardware development based on model architecture and inference costs. other inference providers will try to do the same if they can. https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s
- paxys 2mo agoThat’s ridiculous. Every major AI lab is compute constrained. That’s exactly why nvidia is worth trillions today. If OpenAI had a single extra GPU they’d be using it to run another training cycle or experiment for their next model.
- captainbland 2mo agoTo be fair we don't really know in terms of prices what's real and what's just investor subsidised attempts at market capture at this point. It could well be OpenAI's attempt to drown Anthropic while they've got the halo product if they feel they've got deeper pockets.
- platinumrad 2mo agoWe can guess based on the decisions of other inference providers who serve these models.
- handfuloflight 2mo agoDo you mean if other providers will cut their prices in turn?
- platinumrad 2mo agoYes. For example, third-party inference providers serve DeepSeek V4 Flash just as cheaply as DeepSeek themselves, if not even more so. This is very strong evidence that the low price of the model is not subsidized.
- hzbdhdjs 2mo ago[dead]
- w29UiIm2Xz 2mo agoEnterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
- FuriouslyAdrift 2mo agoPartnerships then consolidation comes next...
- minraws 2mo agoI wouldn't be surprised if they still had some margins since cheaper models are much harder to nail the accurate sizes off, and you still pay 2x for 1M context window. But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300? Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that. Could still be a healthy 10-30% margin. Especially with Terra.
- gentlewater 2mo agoThis is gonna put Sonnet 5 in a really awkward spot.
- bakugo 2mo agoSonnet and Haiku were already in an awkward spot, likely by design. Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it. Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.
- supern0va 2mo ago>figure out what the percentage split between Opus/Fable/Sonnet is. This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.
- petesergeant 2mo agoOpus 5 is not strong enough as the top-of-stack model, and feels idiotic after a week or two of heavy Fable usage, to the point where I'm paying for Usage Credits to keep using Fable rather than having to slum it with Opus.
- StilesCrisis 2mo agoWhen I'm paying for it, Sonnet. When work is paying, Opus 5, then Fable if Opus gets confused.
- Zarathruster 2mo agoI've found that Claude nearly always claims to be Opus in the commit message, regardless of the actual model making the commit.
- baq 2mo ago
- foobar_______ 2mo agoHard to believe numbers. I don't mean that as a critique, but literally I am so impressed. Even if the model is a few percent lower for performance but is 80+% cheaper than competitors and is a US company hosted on US based hyperscaler clouds this is kind of a no brainer. Hard for most businesses to justify otherwise.
- rpdillon 2mo agoThis is exactly the model that DeepSeek V4 Flash followed, and it's been insanely successful as a result, even though it's not frontier.
- ignoramous 2mo agoDeepSeek v4 Pro & MiMo v2.5 Pro (Opus 4.6 quality models for code) are insanely cheap for agent-driven work due to their super low cached-input prices ($0.0036/mtok) [0]. For Luna, the cached-input price drop isn't disclosed in TFA, but the pricing page puts it at $0.02/mtok, & that's 5x more expensive. [0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).
- computerex 2mo agoAbsolutely. Although DeepSeek started announcing "Peak valley" pricing which started making me nervous. I have spent $50 usd in July on deepseek and for that much spend I got SO MUCH mileage. I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.
- ismailmaj 2mo agoit's 80% less cost, not 80% in efficiency gains, could be that Luna was overpriced to begin with, we don't have much info on the models themselves. Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?
- dannyw 2mo agoBeen using OpenAI models since ada/babbage/curie/davinci and at least from my own experience, their APIs feel the same. If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.
- axus 2mo agoSomething can be overpriced and still lose money.
- heisgone 2mo agoLet's suppose each models was subsidized at 70%, so that we only pay 30% of the cost. They would loose much more money per token on the more powerful models. It's in their interest to encourage the use of the less expensive models. Let's say they increase Luna subsidies at 90%. They would still "save" relative to the use of the more expensive models.
- anthonypasq 2mo ago> Let's suppose each models was subsidized at 70%, so that we only pay 30% of the cost. why on earth would you suppose that?
- camel-cdr 2mo agothis type of thing usually means you are the product
- mediaman 2mo agoI don't see how this follows. The cost of nails has fallen by 95% over the last century. It's because the cost of manufacturing has fallen. Not because they are selling the information of nail consumers. Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.
- robocat 2mo ago> The cost of nails has fallen by 95% over the last century No it hasn't! A century ago, some nails cost 2.5% of disposable income, and now the same nails cost 2.3% - only a little cheaper. The cost of nails has remained remarkably consistent for a century. The problem is that you have ignored the depreciation of money. Let's assume California prices and income and pick a bigger retail package of nails as you might use for building a house. The numbers used to calculate percentages: in 1926 a 50lb keg of 4" nails was $2.75 and median after tax income might be $108 per month. In 2026 a 50lb carton of 4" nails is $106 and income might be $4,516. Albeit I assume nails are now more readily available and the quality of nails is likely better; and perhaps I should have compared galvinised nail prices.
- eli 2mo agoThese are the API rates. They don’t train on prompts from the API. In what sense are you the product?
- 827a 2mo agoVera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
- cousinbryce 2mo agoIn a data center that is power constrained but not space constrained they could build out new racks and flip the power from the old racks. Wonder if this will lead to moderately used server GPUs on the secondary market someday.
- Yopolo 2mo agoI don't think they overengineered a DC like this. Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too. It will be swooped of the market the second it hits the market.
- FuriouslyAdrift 2mo agoAMD MI400 series is already shipping to customers (basically everybody) and it is crazy fast (8x to 10x faster than the previous gen and beats published Vera numbers in FP8, loses in FP4) and 432 GB per chip. 72 chip unified rack architecture (Helios) already shipping and projected to also beat Vera in NVL72. MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2027 prob).
- blovescoffee 2mo agoAbsolutely true although at some point it's not just raw numbers but also the kernels that run matmuls and there seems (from outsider perspective) to have been more optimization in the cuda kernels
- jpadkins 2mo agoWhen model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier. The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.
- jrflo 2mo agoBurning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real
- jaggederest 2mo agohttps://taalas.com/ https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second. https://chatjimmy.ai/ https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
- iamjackg 2mo agoHoly crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?" I pressed Enter, and the response was instant. > Generated in 0.037s • 14,205 tok/s This is unbelievable.
- HDThoreaun 2mo agoFor what its worth the frontier lab models can surely be a lot faster if they wanted them to be but theyre supply constrained so theyre doing stuff like multi tenancy. Since you cant self host them no one outside the labs really knows speed as a solo tenant
- WarmWash 2mo agoTotally possible that humans aren't actually that intelligent.
- ceroxylon 2mo agoAs well as the existing intelligence being swayed by emotions, hormones, circadian rhythms, stress, peer pressure, propaganda, and survival instincts.
- subw00f 2mo agoWhy does it matter? This is completely based on data produced by humans.
- customguy 2mo agoThat's a bit like saying a tail is swayed by a dog, as if it could exist without one, or would have anything to do if it did.
- afry1 2mo agoAs if ALL OF THAT doesn't represent inherent and crucial elements of judgement, and therefore"intelligence" itself. We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works. All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was. "Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.
- a13o 2mo agoIMO we won’t see AGI until the models are sitting inside bodies with senses and gut biomes and opposable thumbs and the like, all influencing the context window in real time. Intelligence is a full-bodied experience
- re-thc 2mo ago> Seeing spikes like this makes me question about where the floor really is. You mean they increased the price and then cut it back and now it is amazing? Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park. Not that this isn't good news, but what's impressive?
- zzleeper 2mo agoHad to ctrl+f for someone saying this. I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.
- aesthesia 2mo agoLuna's now cheaper than 5.4-nano (for output tokens). That's a significant improvement.
- paytonjjones 2mo agoAccording to https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/, Luna at Max at it's previous cost was comparable in both performance and cost to Sol at High. With an 80% reduction in cost that becomes a ridiculous outlier in efficiency.
- solarkraft 2mo agoThey have no (other) equivalent to nano, so it makes sense that it’s much cheaper now. It may have been better, but it was also hell of a lot more expensive.
- visiondude 2mo agothere is a ton of downward price pressure from Chinese open weight models
- Yopolo 2mo ago5-10% over months would still be quite crazy. But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google. Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.
- arjunchint 2mo agomore like they were facing pressure from chinese models, and dropped prices and now their margins are squeezed
- Der_Einzige 2mo agoUntil I stop getting downvoted for asserting that these guys are profitable per token, HN is going to continue being pikau face shocked at easily predictable things that any serious AI researcher would tell you, and has been telling you for years now!!!
- SwellJoe 2mo agoI think there's a ton of room for efficiency improvements in how the models are built and run, and I think OpenAI has both prioritized that work and figured out a lot of the tactics (and borrowed some from the Chinese models like DeepSeek and Kimi, which have published a lot of their research and tactics for running big models fast on minimal hardware). I think there's also a new generation of hardware in the past year or so tuned specifically for LLM workloads, where it was almost an accident that GPUs worked to run LLMs before. So, while there's still this ridiculous shortage of hardware, what is being delivered is much faster and cheaper to run for these specific workloads. I wasn't expecting it to happen from the US vendors, though, as they've spent so much capital to get to where they are they need to make huge margins on inference to pay it all back. I expected the Chinese models who're running much leaner operations to be the "frontier" on costs (and they have been). But, I'm glad to see OpenAI joining the "cheap and cheerful" models party. There's a lot of work in that area of capability. Probably most work people are doing falls into that area of capability.
- onlyrealcuzzo 2mo agoPrices have consistently gone down 90% every 18 months like clockwork for about 5 years for the same level of quality. There is no end in sight for at least another generation. There is ZERO reason to believe models 1/10th the size of frontier are completely capped on intelligence and impossible to get smarter. They have consistently compressed the intelligence of larger models. You'll see it first on the small end, when they stop being able to compress intelligence, you know that will slowly bubble up and up the chain to larger and larger models. There's no evidence we've reached that at the bottom.
- gameshot911 2mo agoIn fact there's very good reason to think intelligence has much more room to be compressible - the human brain, for one.