6 ms·
Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a m
by unrvl22 4mo ago
Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official ZAI api rates which are already like 10x cheaper than Opus. (Crof and Umans btw)
This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API prices and speeds mean nothing for open source models.
- cedws 4mo agoIn my org everyone is extremely Claude-pilled to the point you’d think it’s the only LLM that exists, purely because it caters to non-engineers within enterprises.
- unrvl22 4mo agoI cancelled my claude sub after realizing I can burn 300m tokens a day of this quality, for $50 a month.
- spelk 4mo agoWhich coding plan are you using? How are you finding it?
- Hamuko 4mo agoI’m not that interested in models that I can’t run on my desktop for ~0€, which is my AI budget.
- igravious 4mo agoCool beans. You're not the target audience then.
- Hamuko 4mo agoDid I claim I was? I just said why I and people like me are not talking about it.
- simianwords 4mo agoand he said its cool
- andai 4mo agoElectricity cost seems to be about $30/month for a 32B model on a GPU. It's probably better on Apple hardware. https://github.com/QuantiusBenignus/Zshelf/discussions/2 https://github.com/QuantiusBenignus/Zshelf/discussions/2 Not accounting for hardware, of course :)
- Hamuko 4mo agoMy Mac Studio uses about 60–80 watts whenever I’m running a model (as measured by the system metrics), so it’s less than 2 kWh/day at full blast. Electricity is like 0.125 €/kWh, so that 24-hour period would be <0.25 €. Not accounting hardware in my costs, since I didn’t buy my hardware for running models. Running models is just something it can do in addition to what I got it for.
- NorwegianDude 4mo agoThe price, processed tokens, and output can be anything, it just depends on what GPU it is. Nvidia GPUs are much more efficient than Apple hardware for inference(and training).
- CuriouslyC 4mo agoBe careful about unofficial providers, a lot of them misconfigure models or stealth quantize them. For a while the difference between Kimi on the official API and most third party providers was 20-40%.
- unrvl22 4mo agothe 2 I mentioned both have a fairly large following, who run benchmarks and absolutely will spot issues.
- cedws 4mo agoOpenRouter should be penalising or banning for this.
- kilroy123 4mo agoThis is my biggest complaint about OpenRouter and I'm a fan. Might be pretty tough at scale?
- alecco 4mo agoWould that align with their VC-backed incentives?
- mrngld 4mo agoIf your users can't trust your product then I'd say that'd be a pretty strong incentive?
- orbital-decay 4mo agoThey have an "exacto" category with providers they supposedly verified
- ComputerGuru 4mo agoThat’s only for tool use.
- embedding-shape 4mo ago> Why aren't more people talking about this? Wasn't this released like 2 days ago? Everyone is still evaluating and playing around with it, things like the submission is just starting to come out. Give it some days at least before jumping to conclusions, ideally weeks.
- Schiendelman 4mo agoTo answer the question in your first sentence - because it's VERY computationally (ha) expensive as a human being to keep up with all the options. It's also very hard to figure out how to run a model like this. There's no installer. If you really really care, which 99% of people do not, you have to google a guide, and then find out it's out of date... I've tried a number of these, and the learning curve is very steep compared to "install Claude Code and pay $100/mo". There is no way saving me $50/month matters compared to figuring that out.
- andai 4mo agoBut it just works with Claude Code? They have a guide on their website. https://docs.z.ai/devpack/tool/claude https://docs.z.ai/devpack/tool/claude Here's my setup. I add this to my .bashrc export ZAI_API_KEY="your_key_here" alias claudez='ANTHROPIC_AUTH_TOKEN="$ZAI_API_KEY" ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic https://api.z.ai/api/anthropic" ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5.2[1m]" ANTHROPIC_DEFAULT_SONNET_MODEL="glm-4.7" ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-4.7" claude' Then I just run claudez pro tip the same thing works with deepseek https://api-docs.deepseek.com/guides/anthropic_api https://api-docs.deepseek.com/guides/anthropic_api Even more pro tip: Claude Code can set this up for you haha
- Schiendelman 4mo agoSure, I'm not saying I, a software engineer, cannot do this. I'm saying it's significant onboarding friction. Unless this were a massive differentiator, people aren't going to be "talking about it" the way GP suggests!
- fc417fc802 4mo agoYou're seriously suggesting that setting up opencode or tweaking your claude code config or etc is too much trouble to be worth saving $50 /mo? That's absurd. Doubly so when the audience in question is already using LLMs so ... just ask your existing LLM for help if it seems daunting.
- stanac 4mo ago> Some are even offering API rates at 3x lower than the official ZAI api rates Looking at openrouter [1], some of the cheaper offerings are for quantized models. Not sure how much intelligence is lost in quantization. And they are not 3 times cheaper. Where did you find 3x lower prices for APIs? I am considering skipping open router and using them directly for that price. edit: I see, croft [2] 8bit for $0.50/$0.08/$2.20 [1]: https://openrouter.ai/z-ai/glm-5.2 https://openrouter.ai/z-ai/glm-5.2 [2]: https://ai.nahcrof.com/pricing https://ai.nahcrof.com/pricing
- benjiro29 4mo agoNeuralwatt ... When you reverse calculate the actual energy usage / price on a token basis, the gap is large. I do not have GLM 5.2 numbers because the whole default max setting is overkill. But GLM 5.1 numbers had it at 12x cheaper then API rates. And about 2.5x more tokens vs zai their own subscription service. Yes, its FP8 but lets be honest, do we know for sure that even zai runs at FP16? I learned a long time ago with Claude and Codex how much cheating happens on model levels, even from the big boys.
- spelk 4mo agoPlease correct me if you have contradicting data but: Neuralwatt's price per token vs price for energy comparison doesn't seem to take into account the cost savings from cache hits that other providers offer on pure token rates. The comparison seems to assume every input token is a cache miss. On top of that, the cloud offering doesn't seem that well-run, they randomly blocked a colleague's API key for a couple days without any heads up, had a weird rate limiting bug and they have been deprecating models without redirects with very short notice, all while taking weeks to onboard new models. I assume some of these problems would be addressed if we had an SLA/enterprise contract. It's a promising idea though. They offer a $5 trial credit (with an aggressive rate limit) though so no harm in trying it out.
- benjiro29 4mo ago> doesn't seem to take into account the cost savings from cache hits Absolute false information. From my usage panel for this month: * Total Tokens 1.1B * Cached Tokens 1.0B 97% of prompt tokens * Cost energy pricing $26.58 The energy pricing is higher then what i actually pay because its a mix of token billing and partial subscription (60% extra "power"). From the $50 subscription, i have about 3/4 left (4.21 of 16.0 kWh used this billing cycle). Used $5.5 in token billing. That was running 82.0% GLM 5.1, and 18% GLM 5.2. Yes, i have been busy ;) My actual usage if we look in dollar value was ~ $18. For your information, that is cheaper the MiMo v2.5 Pro from Xiaomi as there i was doing around 450.000t per cent. And they have the same 75% cheaper prices like DeepSeek. MiMo has a issue with cache retention between session prompts what hurts them vs DeepSeek. Yes, DeepSeek v4 Pro is 2.5x cheaper but nowhere near GLM 5.1, and especially not GLM 5.2. In case your wondering, zai subscription light is about 80m token / week limit. So on a token/cent price, neutralwatt is about 3x cheaper (and not 5h, week limits to maximize/frustrate). > all while taking weeks to onboard new models. Took them 1 day to include GLM 5.2 ... Yes, the remove old models fast because they do not have the server capacity to keep old models around. > I assume some of these problems would be addressed if we had an SLA/enterprise contract. Its a small team, not a big huge company. From my experience so far, seen a 2 timeouts, and sometimes slow speeds as servers get overloaded. For what i am paying for GLM ~5.1~ 5.2 ...
- anuramat 4mo ago> unlimited tokens for $50 a month link? > Why imho everything but opus produces unusable code (fable was even better...), eg gpt5.5 seems to write the absolute worst code that still technically solves the problem; tbh I'd be totally willing to trade "raw intelligence" for "code taste" more labs need to figure out whatever anthropic did to destroy everybody else on frontiercode bench
- CuriouslyC 4mo agoOpus has the nickname "Slopus" in a lot of circles for a reason. It can write nice code in isolation, but the way it organizes that code and its rigor in addressing edge cases/making sure things are robust leave a lot to be desired. Opus is particularly famous for having a real problem reinventing stuff that already existed in the codebase because it wanted to get to work before exploring sufficiently.
- anuramat 4mo agowhat you're describing doesn't sound like such a big deal -- it's (A) obvious during review, (B) easy to fix in a single prompt, (C) simple enough to fix manually, (D) can be mitigated with tokenmaxxing (agent review passes, prompting, subagents, etc) regarding edge cases -- less is more in my experience, as removing is harder than adding
- knollimar 4mo agoIsn't it closer to sonnet?
- redox99 4mo agoDefinitely opus level for coding.
- smith7018 4mo agoDo you have benchmarks or at least anecdotes to back that up? I'm not arguing with you; I would just love to see some proof that open models are getting as good as Anthropic's models.
- unrvl22 4mo agolook at benchmarks, use the model yourself. Im usually first to call BS on every chinese model that says they are as good as Opus. this is finally the first one that actually is. It is a massive jump from every other previous chinese model.
- smith7018 4mo ago"use the model yourself" I wish I had the time to set it up and work on side projects but unfortunately life and work have been crazy (as I'm sure many here feel). That's why I asked for anecdotes about it.
- redox99 4mo agoI've been running some test prompts comparing frontier models for webdev, particularly pretty visualizations, physics / orbital simulations, etc. Do note that GLM is not multi modal, which can be a deal breaker. And these open models are not good outside coding.
- knollimar 4mo agoOic I misremembered OAI scores, I thought Sonnet had 51
- sinatra 4mo agoI've tried Chinese open models few times before. They were fine, but they didn't come close to the benchmarks they were claiming. Now, maybe GLM 5.2 is close to Opus 4.7, but I don't wanna keep checking them and keep finding that they're still benchmaxing and aren't at GPT (my choice) or Opus level. The boy who cried wolf, I guess.
- enraged_camel 4mo agoYes, my experience has been the same as yours. I find that the performance of open models is quite acceptable, even good, at one-off questions or small tasks. But they are quite unreliable at long horizon goals.
- shostack 4mo agoWhich of those providers are: 1. Keeping your data private on in the US 2. Not training on it 3. Not quantizing the model 4. Offer reasonable latency adds rate limits
- SyneRyder 4mo agoOpenRouter has a list of providers, looks like NovitaAI would meet those criteria. Though not for $50/mth for 80/M tokens, which I assume is the Z.ai subscription pricing. https://openrouter.ai/z-ai/glm-5.2 https://openrouter.ai/z-ai/glm-5.2 https://novita.ai/models/model-detail/zai-org-glm-5.2 https://novita.ai/models/model-detail/zai-org-glm-5.2
- pranavj 4mo ago[flagged]