10 ms·
I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily progr
by pimeys 3mo ago
I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars...
This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and OpenClaw just felt... not what I wanted. Two days later and 20 dollars poorer I have what I need: a multimodal agent written in rust that has access to my homelab.
Nothing felt off with GLM. It did what I wanted, was fast, had a decent not very annoying personality and was much cheaper than Opus or GPT.
I used it unquantized through Fireworks, but there are multiple other providers too.
- HKCM852 3mo agoWhich harness did u use?
- pimeys 3mo agoOpencode and Zed about 40/60.
- noncoml 3mo ago[flagged]
- sertsa 3mo agoIts an editor: https://zed.dev/ https://zed.dev/
- HAL3000 3mo agoJust FYI, this question was a quote from Pulp Fiction, the other commenter (mdre) replied also with a quote, that was an answer to this question in the movie.
- mdre 3mo ago[flagged]
- term333 3mo agoPlease take comments like this back to reddit.
- dist-epoch 3mo ago$20 on API pricing or on subscription?
- pimeys 3mo agoAPI, pay per token.
- shostack 3mo agoIf you're using Matrix, consider Hermes as a harness if you haven't already. Native gateway support. I've been primarily using mine through Element and it has largely been great.
- pimeys 3mo agoOh interesting. I basically chose Matrix because setting anything up with Whatsapp or signal was kind of painful and telegram doesn't make it easy to use encryption with bots. I kind of wanted to see if I can make a Matrix agent from scratch with Rust with GLM and it was surprisingly easy. Just make something for myself how I want it. Maybe I'll take a look on Hermes later...
- Barbing 3mo agoVery interesting—Element X solved a lot of the pains of Element (iOS), could be a good solution!
- KaoruAoiShiho 3mo agoAre you sure fireworks is unquant? It's not listing precision on openrouter like everyone else.
- Aditya_Garg 3mo agoIm really curious about this. Why pay API pricing? I burn 1000s of dollars a month of api according to claude usage but only pay the $100 subscription
- SV_BubbleTime 3mo agoThere is a whole iceberg topic on subsidizing. So your question is really “if they’re giving free usage, why not take advantage of it?” I do, so I don’t know the reasons not to, other than to experiment.
- horsawlarway 3mo agoMy increasing frustration with these plans is the harness lock in. Anthropic won't even let you run "claude -p [prompt]" any more... They bill it at api rates. So if you're trying to automate the ai (and seriously, that's the point) the subsidized plans are crippled.
- weird-eye-issue 3mo agoI think they rolled that back
- sroerick 3mo agoI'm using synthetic.new and Neuralwatt with pi and its good and also cheap
- computerex 3mo agoI have had bad experience with neuralwatt GLM 5.2. Seems like they may be using quantized version of the model.
- scottcha 3mo agoHi I'm the CTO of neuralwatt, would love to hear your feedback on what your experience was. Feel free to email me scott@neuralwatt.com. Also for GLM5.2 we run the FP8 quantization at 1M context which is a common deployment target.
- playorizaya 3mo ago[flagged]
- dom96 3mo agoTwenty dollars? How are you comfortable spending that much to write something as simple as a matrix bot? Are people doing this kind of thing just super rich or am I missing something?
- ygjb 3mo agoIt's pretty simple. There are things that I do because it's fun, like gamedev. I hand code that, and don't use LLM tools because I like learning and building. I do lots of utility stuff coding for my wife's business, most of that is stuff I could do in a few hours. It's worth $20 to not spend a few hours doing it. It's a cost benefit tradeoff. I won't learn much fixing WordPress themes or adding a feature to her web page, or setting up an automation for her, so I don't see the point of doing that. Same thing for stuff at work. Oh, the tables/schema changed and my queries broke? I could dork around with spark and cypher for an hour, or I can tell claude to update the queries for the new schema. At the rate I am paid, spending on Claude tokens is generally a better use of my resources. Building a net new solution? Coding tools take a back seat until I get the core logic right, then I let automation handle web page and UI scaffolding.
- adamtaylor_13 3mo agoIs spending $20 considered "super rich"?
- yard2010 3mo agoRecall that the marginal utility of money diminishes when you have more of it - when you have a lot of money it's easier to turn it into even more money, and vice-verca. It's not linear. So 20$ difference has exponential not linear influence on "being rich".
- copperx 3mo ago$20 is really cheap for the amount of work saved, considering you're in the US.
- annzabelle 3mo agoA lot of people spend $20 on a hobby for an hour of enjoyment a couple times a week. Not odd at all to do that for a few hours of coding if you find it fun. It could be a day pass at a bouldering gym or a yoga class or amortized running shoes/garmin/electrolytes.
- gertlabs 3mo agoGLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 is just shy of Opus 4.6 on average. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings But when factoring in performance/cost, GLM 5.2 is the frontier model.
- jchw 3mo agoAfter having used GLM 5.2 and Opus 4.8 for enough time, I'm very unconvinced of the benchmark maxxing claims - if anything, GLM 5.2's rather lackluster performance on benchmarks compared to Opus 4.8 paints the opposite picture when compared to the subjective experience. When I first used Opus 4.8, I threw several different workloads I had at it - I have Claude doing a lot of misc projects whose primary purpose is pretty much just studying what AI agents can do for my own curiosity and no other reason. Opus 4.8 was one of the first models I ever snuck in there that basically ran out of control. No previous Opus or Sonnet model I had used ever did this. Within hours every agent I had running was writing non-sense tool calls that echoed pretend commands that didn't exist, like 10 in a row, and talking about the "tool channel" being dirty. I switched back to Opus 4.7 and assumed Opus 4.8 was legitimately just broken. I did come back to Opus 4.8 and found that it was indeed, pretty powerful. But that initial experience has stuck with me on just how narrow of a perspective any given test or benchmark is guaranteed to have. LLMs are too broad, it really doesn't matter what you try to do in your benchmark, you will necessarily get a limited view of what the model is capable of and its shortcomings. This will remain true for at least as long as models are susceptible to massive swings in performance based on randomness and minor differences in prompts and other environmental factors. I'm not saying benchmarks are useless or that your benchmarks are not possibly closer to the truth either. All evidence at least points to the idea that Chinese models perform very well in coding but often have more mixed results on other tasks. I'm just saying that at this point, benchmarks feel like they have limited connection to my actual real experiences. GLM 5.2 actually scored kinda meh on a lot of benchmarks (compared to closed frontier models) but my actual experience using it does not match this. And I'm definitely not saying GLM 5.2 is better than the frontier LLMs here, just that the race is close. I still prefer GPT 5.5 right now for code review, I think, and Opus clearly has some advantages depending on the task. It's just no longer a given that Opus 4.8 will perform better than GLM 5.2 on any given task, so to me the calculus behind "using the best model available" is getting complex and you might need to get a feel for what models have what strengths to really figure it out. I do feel like the "use the best model available" mentality is not going to die any time soon, but if it does die, it will be gradual and start soon for programming. Modern LLMs are still not a full superset of what human programmers can do, but still larger models are definitely starting to hit diminishing returns for tasks at the lower end of complexity, and that is a big deal. It's a weird world where some tasks you can feel kinda confident just throwing Gemma 4 at it and not sweating whether you should use a better model; I've certainly done it for some quick Python scripts or getting an overview of some code I'm unfamiliar with.
- jklmnopqrstuvw 3mo ago> A typical session for me with GPT is usually over a hundred dollars. I don't think a $100 session is "typical". I use GPT for months. $20/m plus plan is enough for my daily work.
- adamtaylor_13 3mo agoIt's really interesting what "normal" is for folks. I use the $200/month Anthropic subscription and use it within a few percentages of my limit every week. I'd blow through $20/month plan in hours.
- jascha_eng 3mo agoShorter sessions more often doing a /clear etc. save a shit ton of tokens. I pay 100 bucks a month but barely use 30% of it most weeks.
- tjwebbnorfolk 3mo agoI have Claude max plan and the vscode claude dashboard plugin has logged about $4k worth of tokens in the past 2 months. I upgraded because I was using my weekly basic plan tokens in like 5 hours. Likewise, I don't understand how anyone survives on the basic plans. It's funny seeing these two camps not understanding what the other is doing :)
- simple10 3mo agoI use an observability tool with claude code [1] that shows me usage including prompt and session cost. Even though I use a max subscription, it's interesting to see what it would cost me if I was using API directly. My typical session ranges from $100-$400 - higher end when using workflows with lots of subagents. $100/session is expected when using the API without the subsidized subscription pricing. Most larger orgs have to use API pricing AFAIK. [1] https://github.com/simple10/agents-observe https://github.com/simple10/agents-observe
- jklmnopqrstuvw 3mo ago>Most larger orgs have to use API pricing AFAIK. There are Business and Enterprise plans, both have discounting.
- TimXare 3mo ago[dead]
- neya 3mo agoI am seeing extremely positive results with Elixir too. Previously I was on Deepseek (deepseek-v4-pro) and GLM5.2 outperforms Deepseek easily. It's been a month since I used any native Claude models (simply because of pricing) but then, GLM5.2 is running for me at $20/day in usage on OpenRouter for GLM5.2. I am not sure if I've misconfigured Claude code or if this is indeed normal usage pricing. But, the output more than makes up for it. However, using Deepseek v4 pro directly from deepseek.com using their discounted pricing is insanely cost efficient. I topped up $10 a month and a half ago and I'm still yet to use up all the money in my account. Here's hoping that SOTA models become even cheaper!
- andai 3mo agoNice. I'm working on an agent too. How are you handling tool calls? I followed this example https://minimal-agent.com/ https://minimal-agent.com/ but I'm running into issues with nested backticks so I'm thinking of making dedicated close tags per tool call.
- try-working 3mo agoHave you tried using DeepSeek V4 Pro instead? It will be cheaper and faster than GLM.
- wahnfrieden 3mo agoWhy are you spending on API for GPT coding instead of stacking 20x subs and using codex-lb?
- pimeys 3mo agoCompany pays API prices so we can use daily the best model for our job without being locked in. Also the team subscriptions started to be more like X per seat + usage...
- wahnfrieden 3mo agoOh it sounded like personal use. I understand the reasons to use team/enterprise accounts, but apart from the policy/management/billing side of it, I still don't understand the value in spending thousands for API instead of hundreds - even when there's argument that one provider is better than another depending on the use case, I don't think that credibly extends much beyond OpenAI + Anthropic frontiers, which both have $200 subs you can stack.
- gguncth 3mo agoWhat makes you use API billing instead of a plan?
- croes 3mo ago> This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Did you program or did you gave the order to an agent to program?
- accrual 3mo agoCould you share more about the homelab project? Is it so you could message your local agent via Matrix and it can poke around the lab, check if services are up, restart VMs, that kind of thing? Would love to hear what you use it for, I'm thinking of building something similar for my lab.
- nullbio 3mo agoWhy use an API when you can use a subscription though? Surely a $200 subscription is cheaper than using GLM 5.2 API?
- frr149 3mo agoHow are you using it? A subscription from z.ai?