5 ms·
I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve b
by leobuskin 2mo ago
I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)!
I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
- api 2mo agoThey're invaluable for developers to fix their code. This is definitely an area where AI decisively beats human devs in a very valuable way. It can try so much surface area so fast. If it won't attack my stuff, it won't help me build my stuff to be secure.
- leobuskin 2mo agoExactly my CoT! I hope z.ai won’t change this behavior after training it on our input the same way as Anthropic did (shame on you, folks, seriously)
- jermaustin1 2mo ago> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.
- leobuskin 2mo agoImpossible with source code, possible to bypass with app/site
- doginasuit 2mo agoDon't we already do this with services like Let's Encrypt, which is arguably more sensitive? If you had the codebase you could fake it, but it would still provide some amount of protection against abuse.
- Someone1234 2mo agoWith Let's Encrypt, all the verification is done on their side with them controlling the connection between themselves and whatever they're trying to verify. In this case, you can put whatever you want between the harness you're running (or modify the harness itself), and essentially "lie" to the model. Any verification technique would be fairly trivial to bypass, while you continue to run the harness locally.
- glub 2mo agoYeah, for web apps, you can trick models by simply proxying it and pointing the models to that localhost. They then think they're not working on a live target. Have personally tested this with Opus and Sol and it works. Classifiers are tricky though. Here's where open weights will win.
- gdhkgdhkvff 2mo agoIsn’t this essentially what anthropic is doing, albeit in a manual fashion? They work with code owners to run mythos and find issues.
- rattlesnakedave 2mo agoYou should try a better harness. Try pi, or ohmypi if you want a good OOB experience
- gigatexal 2mo agoI’m in the Claude code harness for everything boat too. What are the alternatives?
- KronisLV 2mo agoWhat the person above is suggesting: * https://pi.dev/ https://pi.dev/ * https://omp.sh/ https://omp.sh/ (no personal opinions of either, links might be useful) I think that OpenCode is nice, their CLI version is enjoyable and their desktop/web version is okay: * https://opencode.ai/ https://opencode.ai/ I also quite like driving OpenCode through something like Kepler / Paseo and tools like that (with those I can still use my Anthropic Condition by Claude Code being treated similarly - as something that gets tasks dispatched to it, while the GUI I see is Kepler / Paseo). On the desktop side, ZCode was surprisingly usable for something that came out of nowhere (I wasn't aware of it at all before trying out the GLM Coding Plan): https://zcode.z.ai/en https://zcode.z.ai/en
- vadansky 2mo agoLast time I tried some of these, none of them had the "manual mode" that CC has, where it shows you change by change as diffs and you can edit them before accepting and moving on to the next change. I like that because if it's going off pattern I can spot it early on and guide it correctly, instead of having to review the whole completed diff at the end when it's too late. I should spend the weekend checking them out again to see if they added that but I assume with everyone going full agent mode they probably didn't.
- eli 2mo agoThe philosophy with Pi is it is minimal (but functional) out of the box and easily extensible. I'm not familiar with that feature but I would not at all be surprised if someone already coded a Pi extension that does it.
- maayank 2mo ago“ Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.” Cybersecurity capability might be nerfed
- deleted 2mo ago[deleted]
- misiti3780 2mo agohow do you configure claude code to use GLM ?
- leobuskin 2mo agohttps://docs.z.ai/devpack/latest-model#switching-models-in-claude-code https://docs.z.ai/devpack/latest-model#switching-models-in-c...
- deleted 2mo ago[deleted]
- matheusmoreira 2mo agoHow much usage do you get out of it per week? How many millions of tokens? Anthropic was stingy as hell with its Fable and cybersecurity nonsense, switched to OpenAI which is much better but still not enough. I'm tempted to switch again...
- bicepjai 2mo agoYes, I am tired of Claude and GPTs. I am ready to diversify my $300 per month on other vendors. Will try GLM. How was your rate limits and availability experience on $80 dollar plan?
- leobuskin 2mo agoIt’s comparable to Anthropic usage, to be honest. 2x GLM agents ate 18% of weekly usage on this mid-tier plan within ~8 hrs (non-stop work, a lot of tool calls, appx 4 compactions each), I think. I didn’t make a proper statistics snapshot, sorry.
- ma2kx 2mo agoOutside peak hours (which are during Chinese daytime) I dont reach them with a single agent.
- takerofnaps 2mo agoAt my work I have a $500 monthly AI budget. I have been using the $200 Claude subscription and most of my use is with Claude code. I think I'm going to switch to either kimi or glm and use the opencode harness. Both fable 5 and opus 5 have outright refused things like security related bug fixes and making monitoring tools. I am so happy that open models are good now
- gabriel-uribe 2mo agoI generally use the $100-200 Codex/Claude subs, and have been blown away by the usage I get from OpenCode Go at $10/mo. At a minimum, excellent for automatically piping reviews to from Codex/Claude.
- LaurensBER 2mo agoI run my OpenClaw on whatever is the latest GLM model and ever since the release of GLM 5 it has been a smooth ride. The models solve whatever problem I throw at them and the code is good enough that I barely ever have to look at it (to guide the mode). The 5.3 release seems particularly strong, I asked if to audit all the scripts that the previous versions have written and it identified some issues and hard to find bugs. At work, as an experiment, I used GPT 5.6 Luna + Deepseek 4 Flash for a week (I have an unlimited, "within reason", budget at work so normally I just use Fable and Sol) and it's been perfectly fine. These models take a bit longer (more turns) to solve problems so they feel a bit slower but the end result is often just as good or nearly as good. Because they're so cheap you can easily run multiple sessions in parallel so it doesn't really matter that they're slower. I've done a few experiments where I've split my terminal in 4, launched 4 clients (each with a different model, including Fable and GPT 5.6 Sol) and compared the output. For simple and medium complexity work open-weight models are incredible effective. I can highly recommend the 10 USD/month OpenCode Go subscription. It offers pretty amazing value for the money and is a great way to experiment.
- jcomp 2mo agoMy beef with the Fable refusals is that it seems to just be flagging keywords, and also seems it flags on keywords the model itself introduced to the context. In a normal Chat with Fable, something like "How can I exfiltrate a guy from a sticky situation?" reliably downgrades, leading me to believe that Fable just outright refuses once it sees the word "exfiltrate". When it writes a service and names it CloudExfiltrator, the next turn downgrades to Opus. Opus doesn't appear to refuse on simple keywords, but it does seem like Fable's reasoning introduces enough nefarious-sounding context that Opus will then refuse, and I'm stuck playing the new session game despite having done everything correctly myself and having a totally innocuous prompt. At one point, Opus was happy to continue while outputting commands for me to execute on its behalf, but flatly refused to execute them itself through multiple new sessions. To its credit, it openly acknowledged how ridiculous that was and was apologetic for the safeguard. I'm open to the idea of some kind of guardrails, but if Fable is so dangerously intelligent as to require the guardrails you'd think they could come up with something a little more nuanced than a list of bad words. As far as I can tell, they've also not done anything towards improving the situation since the model was released, despite the "deliver more capabilities faster" claim.
- deleted 2mo ago[deleted]
- clbrmbr 2mo agoHow are you using it in Claude Code? What is the native harness that GLM was post-trained in?
- leobuskin 2mo agoThat’s the first model that fits CC as it’s own, zero issues, but probably ZCode or whatever z.ai’s cli is.
- uejfiweun 2mo agoNo no no, absolutely not. There's only one man that should be allowed these privileges, and his name is Dario. Dario alone can deliver us to salvation. The lord himself shalt smite these companies and models from this barren earth, and Dario will rise from the ashes to ascend to godhood. Dario. He alone has the power to decide what capabilities us mere mortals have access to.