4 ms·
I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't tha
by embedding-shape 18d ago
I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.
They must have hit really hard scaling limits if the prices were hiked so much so quickly.
- broodbucket 18d agoYeah it went from a great deal to unviable compared to other providers imo. They really need to find a healthy middle ground
- lompad 18d agoIt just gives a taste of what we are all going to have to pay soon, once the model providers actually have to make money. And the era of "let's charge a dollar for every 10 dollars running the infra actually costs" is rapidly coming to an end. And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.
- chobbledotcom 18d agoThis isn't true, you can pay for GLM 5.3 from a provider like Neuralwatt or Friendli who have no incentive to subsidize or loss-lead their inference APIs
- jdiff 18d agoThis introduces other incentives to cut corners and over-quantize.
- breakingcups 18d agoThey didn't pay for training
- pyrophane 18d agoWhat provider are you using currently?
- bbor 18d agoIt's hard to know, since no one advertises the actual token limits (partially cause they're prolly complex / adaptive). So it seems much more likely that they just offer different pricing tiers than you're used to. Like, the $80 plan is still ~$80 of subscription quota, regardless of what else is offered. For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.
- Daviey 18d agoI paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.
- disiplus 18d agoI also have a legacy pro plan and the only limitation is if you are trying to work in the morning from Europe because you are in the 3x usage overlapping China time but after 12 or so you basically can run it at least for me at least 3 parallel sessions all the time.
- world2vec 18d ago1 billion tokens a day?!! I've done a lot of work these past 2 weeks with GLM-5.3. Like, a lot. And I've just passed 300 million tokens in total. Can I ask where are you using all those tokens?
- wartywhoa23 18d agoSomething like this I guess: https://youtu.be/U-Rqv9dOB1U https://youtu.be/U-Rqv9dOB1U
- tokai 18d ago300M for two weeks is surprisingly low. What are you doing that need so few tokens?
- 18d ago
- asp_hornet 18d agoThe way I look at it, their coding plan doesn’t retain data or use it for training making it one of the cheaper plans for me. https://docs.z.ai/legal-agreement/privacy-policy https://docs.z.ai/legal-agreement/privacy-policy
- andy_ppp 18d agoYou believe any of these companies care about the law? They care about winning and building the self improving AI as quickly as possible.
- asp_hornet 18d agoI too am sceptical but I’ll take my chances. At least it’s helping the open weights.
- criley2 18d agoI believe the that the companies who claim to not train on my data are more likely to not train on my data than the companies who refuse to even claim they won't. Also why Meta gets a +1, just charge less money on the training path.
- orf 18d agoI’m not sure that follows. You’re assuming that all those claims have the same weight, without considering the size, jurisdiction, reputation or even the general vibe of the company making that claim. If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying. These are not equal.
- asp_hornet 18d ago> I’m not sure that follows To be fair, none of us are sure of anything and I think that’s the part that’s most irritating
- Havoc 18d ago>I was gonna ask how people found their coding plans Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate. >They must have hit really hard scaling limits if the prices were hiked so much so quickly. Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.
- probst 18d agoWay to restrictive in terms of tokens provided. I am on their largest plan, and quickly run into their limits. And that is using it selectively in addition to codex.
- _aavaa_ 18d agoTheir plans are still worth it if you use their models. You can see how many tokens you can except to get based on plan here: https://docs.z.ai/devpack/overview#estimated-token-allowance https://docs.z.ai/devpack/overview#estimated-token-allowance The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate). Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow. They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed. They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).
- Schlagbohrer 18d agoThat table assumes cache hit rate of 95% or better. Am I understanding this correctly that people really are doing such repetitive prompts (compared to each other, across the concurrent user base at that time) that only 5% or less need actually be computed by the intended LLM? That is shocking. Is it per-token I wonder?
- workbreak 18d agoEvery tool call is essentially entire prompt so far sent again with the response and that's why cache rates are so high for agentic workloads. This really bites when using expensive models since most models are 1/10 for cached input.
- _aavaa_ 18d agoIf you are using their coding plan for coding, then yes you can easily hit such cache rates, with a good harness. I’m getting 97%.
- bleonard 17d agoWe have been running a lot of agentic benchmarks with the various loops and tool calls on longer threads - we routinely see 90%+ Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90% [1] https://www.induction.ai/docs/benchmarks/tau3 https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon https://www.induction.ai/docs/benchmarks/toolathlon