4 ms·
GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at
by mikenew 6mo ago
GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all.
Some people seem to agree and some don't, but I think that indicates we're just down to your specific domain and usage patterns rather than the SOTA models being objectively better like they clearly used to be.
- operatingthetan 6mo agoIt seems like people can't even agree which SOTA model is best at any given moment anymore, so yeah I think it's just subjective at this point.
- fwipsy 6mo agoPerhaps not even necessarily subjective, just performance is highly task-dependent and even variable within tasks. People get objectively different experiences, and assume one or another is better, but it's basically random.
- operatingthetan 6mo ago>just performance is highly task-dependent and even variable within tasks. People get objectively different experiences, and assume one or another is better, but it's basically random. You are right that this is not exactly subjectivity, but I think for most people it feels like it. We don't have good benchmarks (imo), we read a lot about other people's experiences, and we have our own. I think certain models are going to be objectively better at certain tasks, it's just our ability to know which currently is impaired.
- easygenes 6mo agoUnless you're looking at something like a pass@100 benchmark, the benchmarks are confounded heavily by a likelihood of a "golden path" retrieval within their capabilities. This is on top of uncertainties like how well your task within a domain maps to the relevant test sets, as well as factors like context fullness and context complexity (heavy list of relevant complex instructions can weigh on capabilities in different ways than e.g. having a history where there's prior unrelated tasks still in context). The best tests are your own custom personal-task-relevant standardized tests (which the best models can't saturate, so aiming for less than 70% pass rate in the best case). All this is to say that most people are not doing the latter and their vibes are heavily confounded to the point of being mostly meaningless.
- mentalgear 6mo agoThis. Plus if you want to even attempt measuring real 'intelligence' you want to run a neuro-symbolic, de-lexicalized benchmark (e.g. DL-ReasonSuite, SoLT, GSM-Symbolic) - which none of the providers releasing new models showcase.
- make3 6mo agoThe pass@100 is such a weird critique angle that is surprisingly mainstream; guess what, no one cares if the correct answer is in the top 100, it needs to be the top 1. A model with a better answer in the top 1 is a better model, full stop.
- hamdingers 6mo agoAnd the subjectivity is bidirectional. People judge models on their outputs, but how you like to prompt has a tremendous impact on those outputs and explains why people have wildly different experiences with the same model.
- ulfw 6mo agoAI is a complete commodity One model can replace another at any given moment in time. It's NOT a winner-takes-all industry and hence none of the lofty valuations make sense. the AI bubble burst will be epic and make us all poorer. Yay
- StilesCrisis 6mo agoStaying power is probably the most important factor, which is why I'm thinking Google eventually takes the crown.
- api 6mo agoThey might be converging somewhat. The ultimate limiting factor is training data. Eventually I think they will converge and then the competition will be on memory and compute efficiency, with the best being the smallest maximally capable model.
- Ladioss 6mo agoSOTA models war is the new console war. But more seriously, I can't help but be amused by how emotionally invested in their AI brand of choice people are getting.
- LoganDark 6mo agoThe value in Claude Code is its harness. I've tried the desktop app and found it was absolutely terrible in comparison. Like, the very nature of it being a separate codebase is already enough to completely throw off its performance compared to the CLI. Nuts.
- deaux 6mo ago> The value in Claude Code is its harness If this was the case then Anthropic would be in a very bad spot. It's not, which is why people got so mad about being forced to use it rather than better third party harnesses. Pi is better than CC as a harness in almost every respect.
- enochthered 6mo agoAnthropic limiting Claude subs to Claude code is what pushed me away in the end because I wanted to keep using Pi.
- strel0k1 6mo agoJust sign up for an AWS account and use the Anthropic models through Bedrock which Pi can use.
- adrianN 6mo agoWhy use tricks to support a company that is hostile to your use case?
- seunosewa 6mo agoAPI costs are really high compared to subs.
- solenoid0937 6mo agoThen you aren't the target market.
- 6mo ago
- abustamam 6mo agoWhat is your workflow? Do you use Cursor or another tool for code Gen?
- mikenew 6mo agoI use Opencode, both directly and through Discord via a little bridge called Kimaki. https://github.com/remorses/kimaki https://github.com/remorses/kimaki
- mettamage 6mo agoHmm Will try it out. Thanks for sharing!
- vidarh 6mo agoI feel like it's Sonnet level for implementation, but not matching up to Opus for planning. But I agree it's close enough that it's worth using heavily. I've not cancelled my Claude Max subscription, but I've added a z.ai subscription...
- alfonsodev 6mo agoMy combo is codex and claude basic subscription for planing the hard tasks (if any) opencode with GLM 5.1 (z.ai coding plan) for the actual coding. opencode is awesome I don't miss cluade or codex cli at all, and the z.ai plan is way more generous in compression. I was lucky to subscribe to z.ai coding plan pro when it costed 30$/month, I was surprised now it costs 70$/month. In case anyone wants to subscribe to z.ai with 10% discount [1] * here is the credit campaign rules * [2] - [1] https://z.ai/subscribe?ic=MW6H74HAZ0 https://z.ai/subscribe?ic=MW6H74HAZ0 - [2] https://docs.z.ai/devpack/credit-campaign-rules https://docs.z.ai/devpack/credit-campaign-rules
- scotty79 6mo agoI had one occasion where GLM 5.1 did about 95% of the implementation that I needed but couldn't progress form there. And Codex (free quota) solved the remaining 5% on the spot. I'm super happy with both. I don't touch anything Anthropic with a 10 foot pole.
- _blk 6mo agoWhat hardware do you run it on? Trying to consider the cost of subscription + API vs new HW..
- DeathArrow 6mo ago>GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all. GLM 5.1 is pretty good but there are some "buts". They hiked the prices 2 times this year. I subscribed to the pro coding plan just before the last hike. At the start of the year, they had only 5 hours quota and no weekly quota. And I hit the weekly quota hard. I can't upgrade the subscription to get a higher weekly quota because they jacked up the prices a lot recently. My $30 subscription costs now $72. Previously was $15. Max was $49,then $80 and now $160.
- _s_a_m_ 5mo agoI used GLM 5.1 and it was bad, I have no clue why people claim it is good