4 ms·
For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen. Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.
by user43928 1mo ago
For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen.
Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.
- rohansood15 1mo agoThis is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.
- jrflo 1mo agoThe luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.
- QwenGlazer9000 1mo agoWas it? Given the timing, I think they A. shat their pants since Deepseek flash just came out with insane pricing before the price hikes, and B. Anthropic is really struggling in model tiers below opus. It was smart for them to cut prices regardless of whether they had 80% efficiency gains or not
- throwaw12 1mo agowhat if it was because of quantization and they haven't released the new benchmarks for it? Anything which changes the model needs new benchmarks I guess to compare with other models, otherwise you can benchmark Fable, and distill it to student model and keep claiming this is the Fable model
- dannyw 1mo agoARC Prize has retested Luna after the discount and validated identical performance. (Also, quantization isn't inherently bad or damaging when done properly, e.g. QAT). These APIs are used heavily by enterprises at scale; with lots of performance telemetry, live evals, etc. You can't really silently nerf API models at scale without people noticing. Of course, what I said doesn't apply to non-API consumer sub models; there's many documented and officially confirmed instances of under-the-hood "juice/effort" adjustments. (Juice = a number your effort tier maps to underneath the hood; much like Inkling's effort=0.00 to 0.99).
- mattalex 1mo agoYou assume that openai's inference is profitable and that they aren't just trying to bolster revenue before their IPO. The only indication that openai is profitable comes from openai (whom I wouldn't trust with any statement, especially when it comes to profitability). In fact there is evidence that inference is not profitable simply because the rate of losses doesn't seem to reduce as revenue increases: if inference had great margins, we would expect that as revenues increase, the amount of spend on training reduces as a fraction of total expenses. Since the loss-making fixed costs shrink as a fraction compared to the profitable inference, we should expect profitability to rise with total revenue. However, all leaks of openai's numbers seem to suggest the opposite: as revenues increase so do the losses.
- hluska 1mo agoI don’t pay OpenAI’s bills - I pay what they charge me. Their cost accounting isn’t relevant to a user.
- aaa_aaa 1mo agoArgument was that open ai cannot be profitable with this. But sure, use it while you can.
- mediaman 1mo agoYou can make the other argument that China subsidizes the price and that they can't be profitable at this pricing level. From an industrial strategy standpoint, they already do this for many other industries with huge subsidized state loans. So we can go round and round on this, each with our made-up objections about how it's temporary or unrealistic or impossible or whatever, or we can just accept the prices as listed and use that to guide our economic decisions.
- aaa_aaa 1mo agoPrivate companies cannot play that game too long. Profit from current state of AI is a mirage and sooner or later stuff will hit the fan.
- Almondsetat 1mo ago>If you compare launch pricing Why?
- nl 1mo agoWhy would anyone car what the launch price is? Comparing launch pricing is just an odd thing to do.
- rohansood15 1mo agoBecause labs can learn to optimize inference post launch, plus can move to use bigger/better clusters depending on demand. It is not impossible to imagine Qwen cuts prices further with QAT/MTP-like improvements.
- nl 1mo agoOr they could move from highly subsidized models like the Deepseek 4 launch pricing. Launch price is just like any other price. It's just a price. It's impossible to guess what might or might not happen. Compare the price now.
- criley2 1mo agoThose prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights models use dramatically more tokens and time to complete tasks than many frontier models. In Artifical Analysis's cost per task, Luna(max) costs $0.05 per task, and Qwen 3.8 27B costs $0.25 per task, a 5X increase. We'll see how 3.8-flash-next does.
- hadlock 1mo agothe important thing is that Qwen 3.7 27B will run unlimited jobs on my consumer grade laptop at 60 tokens/second for free, forever, in about 1-2 years
- villish 1mo agoThats only important if running it locally is critical for privacy reasons or just as a hobby. Time has a cost in business. If a model needs 30 million tokens to achieve a similar result as another that can do it in 10 million, that 60 tokens per second will take a long time.
- hadlock 1mo agoRight now qwen 3.6 35b-a3b has a success rate of 92% and qwen 3.8 27b has a success rate of 96%. But the 35b moe does about 1080 tokens/s at concurrency 54, vs 480 tokens/s at concurrency 28. For our specific workflow on blackwell. Of course enormous batch jobs are different. I was explicit when I said consumer laptop.
- criley2 1mo agoIt's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is interesting in this space. Not to say that there aren't other benefits of running models locally, I loaded Qwen 3.8 27B 6bit MLX just yesterday.
- claudeIsDown 1mo agoSounds like discrete propaganda