5 ms·
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courte
by aliljet 2mo ago
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
- arcanemachiner 2mo ago> this is just GLM 5.2 with post-training magic Isn't post-training turning out to be the most important part?
- HarHarVeryFunny 2mo agoIt basically has been ever since they started using RLVR for reasoning (esp. coding & math), with the DeepSeek-R1 paper being what let the cat out of the bag. The Gemini 3.7 Flash model released yesterday, and all the 3.x Flash models, are still based on the Gemini 3 pre-training run from January 2025 !!
- bertili 2mo agoDwarfStar (https://github.com/antirez/ds4 https://github.com/antirez/ds4) supports GLM 5.2 and DeepSeek. Not only for toying, but for getting work done.
- VulgarExigency 2mo agoSince GLM-5.3 has the same base model as 5.2, DwarfStar should support it as well, once the weights are released, right?
- slopinthebag 2mo agoNope
- kouteiheika 2mo ago> This is absolutely still shy of Sol and Fable Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they can't use it for that anyway.
- bpodgursky 2mo agoI don't understand all this spite about "rich friends" when it was the US government that shut Fable down for not adequately blocking cyber capabilities. I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
- deepllm 2mo ago"Mythos" is the cyber-security equivalent of Fable (without guardrails), and only a very select few corporations have access to it. Fable is their version with guardrails on everything except "Make me a pelican svg" or "create a to-do" app, that is the version that the government banned
- bpodgursky 2mo agoI know all this? Only a few corporations have Mythos because the US government is whitelisting them one at a time. Anthropic releasing Mythos to the public was never on the table, they would have been shut down in milliseconds by the feds if they tried.
- deepllm 2mo agoBefore the US government had anything to do with this, Anthropic were fear mongering Mythos (BTW, Amodei also fear-mongered GPT-2, so this is a normal pattern in their operation) calling it "too dangerous to release", and back then only Anthropic was in charge of the whitelist. Then the government believed Amodei's bullshit and this is a result of that, this was all self-inflicted.
- teravor 2mo agothe difference is that with open models jailbreaking is trivial if you know what you are doing so this makes a frontier open model infinitely more useful for certain tasks seeing as closed frontier models will just refuse (and jailbreaking them is a waste of time when you have good open models). in some cases (mainly reverse engineering) I have observed GLM 5.2 jailbreaking itself with no effort on my part, the thinking trace revealed that it did some mental gymnastics to pretend it was a crackme or capture the flag competition.
- bossyTeacher 2mo ago> This is absolutely still shy of Sol and Fable, but only just by a hair. Even if there was a small/medium gap, the fact that this is a free model beats both of the above on pure economics.
- deepllm 2mo agoRealistically, you're looking at least 2x DGX sparks to run this at a 2 bit quant, but quantization really lobotomizes models so it's just better to run DSv4 flash at full precision. 4x DGX sparks should let you run this at 4 bit at least and there are some folks who ran GLM 5.2 on this configuration in r/LocalLlama
- teruakohatu 2mo agoHow fast are 2x or 4x DGX? I only have one and am wondering what the benefits are of getting another. I feel I will be disappointed…
- deepllm 2mo agoIf you can afford it, another DGX spark is worth it imo. Especially since, owning just one, you have a $1000 ConnectX7 card that's unused. You can find speeds here: https://spark-arena.com/leaderboard https://spark-arena.com/leaderboard
- colingauvin 2mo agoFor DS4 Flash, with 2x Sparks, I am getting 35-85 TPS in single stream, fresh context after quite a bit of RoCe config and the DSpark MTP, on vLLM with Ray and tensor parallel = 2. For multi-stream, it tops out all stream at well north of 100-120. This all degrades with context, but I rarely fill context that much, and if I do it's coding where it's non-real-time. For something like GLM, it's larger, has a larger number of active experts, and doesn't support tensor parallel. This means performance doesn't really scale with more Sparks. You can layer split, but then you are still seeing each layer in series and so if anything performance gets slightly worse. I would not expect more than 10-20 TPS on GLM with 2-4 Sparks.
- disiplus 2mo agoi run flash v4 at 2bit, its pretty great and on my tests against full model It didn't lose any capabilities. It just was thinking more. So you don't have the same efficiency.
- colingauvin 2mo ago
- MangoCoffee 2mo agoOpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize. These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin. I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.
- grey-area 2mo agoIt is impossible to justify the absurd private valuations they have given themselves in collusion with investors. I wish they had tried to IPO because then we’d see the judgement of the market on this. But that’s why they didn’t this year. How long can they keep up the charade that their models are uniquely valuable and on the path to AGI?
- andsoitis 2mo ago> private valuations they have given themselves in collusion with investors. What's the collusion?
- andxor 2mo agoFable finished training 6+ months ago. At this point, Anthropic only needs to release models to the public when the competition forces them to. OpenAI also has a better model (Astra) that they haven't released yet.
- nozzlegear 2mo ago> At this point, Anthropic only needs to release models to the public when the competition forces them to. Assuming the government allows them to lol
- mbil 2mo agoYes it seems like the thread is discounting that frontier providers are likely already baking new, stronger models. I agree that GLM and its ilk are quite good, but having used them I’m not convinced they’re on par with eg Opus in terms of things like tool calling. And they’re fast but less capable so I spend about the same amount of time with them, just with more hand holding. Maybe this is a harness limitation. I know on paper they seem comparable but anecdotally and qualitatively they’re not as useful as the frontiers’, so maybe there’s some truth to benchmaxing claims. For some workloads the distilled models may be good enough, and I suspect at some point there will be diminishing returns to spending a premium on frontier models, but I don’t think we’re there yet. That said I’m continuing to try them. The question is whether this steals enough marketshare from frontier providers that they don’t have the capital to train the next model iteration. The open models are going to push down the unit price of an intelligence-token, but there will still be a market for a smarter bot. And as intelligence gets cheaper, the demand for it will rise (see Hank Green’s Jevons Paradox video). Not to mention there’s all kinds of other directions to go at the frontier (world models, robotics, video gen, etc). Another thing, and this is pure speculation, but if the Chinese model providers already discovered the decrypting COT trick and leveraged it to do RL training, and assuming frontiers plug that hole, then maybe future distillation will be harder.
- justapassenger 2mo agoIt’s not that frontier providers won’t keep on making good/leading models. It’s whether you absolutely need the latest capabilities (at the cost of very high prices, sending your data to them, and being totally at the whim of 2 companies, that can shut you off anytime for any reason). With how good LLMs are already, there’s tons of tasks where not being at the absolute bleeding edge doesn’t matter, especially when you add cost/freedom/supply chain risk/not leaking your data. Even more - there’s increasing number of companies that give you ability to post train open weight model yourself, for your own use case. Given how many of the gains today are from post training, if you post train it for your specific use case, you’re very likely get model that you own, that works for you as good as frontier, at the fraction of the cost. That’s not something for an average Joe to do, but for any bigger business with big spent it’s only natural thing to look into. Just one example - cursor composer - that’s fine tuned kimi. It’s not whether frontier labs will stop releasing models. It’s whether they can generate enough profit out of them. 2 years ago (even 1) they basically had monopoly and combined with demand explosion as capabilities exploded - valuations grew to insane levels. But math now looks different - they no longer have monopoly.
- wren6991 2mo agoThe thing that blows me away is it does this at one quarter the total parameter count of K3 (and 40% active parameter count). There's plenty of room at the bottom. > How are you all toying with running this kind of thing in a mega quantized way locally? Sure, let me answer that in excessive detail. I briefly tried running the UD IQ3_S quant of GLM-5.2, which is 288 GiB of weights (301 GB). Setup was: llama.cpp, 1x NVMe SSD (Evo 980), 64 GiB DDR5-5200, i9-13900HX, and 1x RTX Pro 6000. Token generation around 0.7 t/s. Not remotely usable interactively, but something I could plausibly push a codebase into and come back to a review in a couple of days. There's potential for that hardware to go much faster, but current local inference backends make poor use of the memory hierarchy. Ideally I would have: always-active weights, KV and hot expert cache in VRAM; warm expert victim cache in host RAM; and disk as a last resort. Instead it's 1/3rd of the layers fully pinned in VRAM (all experts), and 2/3rds running wholly on the CPU with mmap()'d weights. The CPU cores spend most of their time sleeping on disk fills. llama.cpp has backed itself into a bit of a corner architecturally by trying to support all models on all possible backends. If you look into how their "MoE offload" feature works (not viable for me because it requires enough host RAM to permanently pin the weights) you very quickly realise it's "oops, all bubbles!" due to the static compute graph splits. There are more focused frameworks like DS4 [1] and Colibri [2] which have better support for streaming weights from disk, and support GLM-5.2. Obviously I wouldn't recommend my setup for huge models like GLM-5.2. Supposedly it can just about be squeezed into 3x GB10, or run comfortably on 4x GB10 (tensor-parallel) for multi-user serving. I'm not sure whether that qualifies as local, but it's at least not a rack. [1] https://github.com/antirez/ds4 https://github.com/antirez/ds4 [2] https://github.com/JustVugg/colibri https://github.com/JustVugg/colibri
- cyanydeez 2mo agoI'm hoping colibri can start pulling in specifically designed models for the heirarchy of decoding. It seems like we should be able to get smarter MoE models that can do the work.
- irthomasthomas 2mo agoHave you seen the news about decrypting the hidden COT in U.S. models? [0] The decoded logs revealed instances where Claude memorized answers to test questions beforehand while making its final output look like it had derived the answer step-by-step—hiding the memorization from the user. 0: https://www.alphaxiv.org/abs/2608.09867?hl=en-GB https://www.alphaxiv.org/abs/2608.09867?hl=en-GB
- xmcqdpt2 2mo agopdf https://arxiv.org/pdf/2608.09867 https://arxiv.org/pdf/2608.09867 for some reason I couldn't find any way to download it from that website.
- r0fl 2mo agoEach time I try to use GLM it is under heavy load and I get downgraded to the older model. So much so that I have given up trying to stop wasting my own time. I rather pay a few bucks more and not have to deal with that nonsense
- erinnh 2mo agoAre you using their web frontend? That seems really broken. I also get downgraded there all the time, but via API all is fine.
- HarHarVeryFunny 2mo ago> This is absolutely still shy of Sol and Fable, but only just by a hair What's crazy is that this is a relatively small model - approx. 750B total, 40B active params, while Sol and Fable are one or two tiers above that (Kimi 3 and Qwen 3.8 also ~3T params).
- lossolo 2mo ago> but this is still just GLM 5.2 with post-training magic. So exactly the same as Opus 5 and GPT 5.6 Sol. It's all "post-training magic".
- segmondy 2mo agoI can run this at home. No guardrails, this is not shy of Sol and Fable, this crushes them in my book. It's not just about evals, but what I can do with the damn model.
- varshar 2mo ago>> This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. Agreed. This release is the first time I'm able to employ a GLM model to write a substantive plan for a complex Clojure PR [1] with both Opus 5 and GPT-5.x playing supporting / reviewer roles. Initial results are __very__ encouraging. GLM 5.3 - - follows directions, - digs into detail, and - correlates well. Still not confident about entrusting GLM with implementation - but IMHO, western labs are entirely cooked. [1] 2K LoC PR in a 55K LoC Clojure + Clojurescript repo
- johnnyApplePRNG 2mo agoThese resets are not nearly enough. I am in the process of creating my own Pi Coding Agent harness to leverage the power of Deepseek V4 Flash 0731 and other models (you can do that when you build your own harness! easily route opinions from other models whenever you're stuck, etc) and cancelling my Codex account next week.
- ShinyLeftPad 2mo agoWhat is reset addiction?
- Sabinus 2mo agoOpenAI, or specifically one guy on Twitter, seems to be so regularly resetting quotas that it's becoming the new normal and it'll suck when it stops happening.