5 ms·
I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not
by XCSme 3mo ago
I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster.
Weak spots (categories it fails):
- Trivia — 0/3 - basically not much built-in knowledge
- Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls
- Puzzle Solving — score 77, flubs carwash-like tests
[0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-medium/anthropic-claude-sonnet-5-medium/anthropic-claude-opus-4-8-medium/z-ai-glm-5-2-medium/ https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
- XCSme 3mo agoAs always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.
- yieldcrv 3mo agoWhat’s everyone favorite GLM provider? z.ai doesnt always have the most reliable AI but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh So I’ll consider all, fastest tokens/sec wins
- eli 3mo agoFireworks.ai is solid. And if you care more about speed than cost they have a "fast" variant that I think just throws more hardware at the model for about 2x the cost.
- david-gpu 3mo agoThe privacy policy indicates that they track you and share your data to ad networks like Meta. Yikes.
- pranaybhatia 3mo agoHi, PM at Fireworks here. We have zero data retention so we do not log any of your API requests. Realize you're talking about website activity which is different and will check and update on that too.
- ricardobeat 3mo agoFast variants are usually quantized to NVFP4, which incurs a slight degradation in intelligence.
- eli 3mo agoIt isn't.
- Onavo 3mo ago> the Chinese one won’t reply to subpoenas so thats a value add tbh That's not something that's definite. They are not quite like the Russians. A lot of the governments in Asia are overly pragmatic and will happily strong arm their companies to throw users under the bus for the sake of a trade deal. There's a reason why Snowden ran to the Russians and not China. Also, if they have any subsidiaries in the US, they may not have a choice in the matter.
- reissbaker 3mo agoI'm biased because I run an inference company, https://synthetic.new https://synthetic.new. That being said I think we're pretty good at serving at GLM-5.2 — and other models, like Kimi K2.7! — and our privacy policy is quite good: zero data retention for prompts and completions on API requests. Our average streaming TPS for GLM-5.2 (aka, tokens after factoring out time-to-first-token, which varies based on geography) is 97tps over the last 24hrs, although it's slightly lower at peak traffic in the mornings PST where it's 50-70 tps. We're also subscription-based which is nicer for coding than e.g. Fireworks which is per-token billing.
- yieldcrv 3mo agogot a 500 error page on the site's chat, but I'll try the API
- reissbaker 3mo agoInteresting: I don't see anything in our error logs but we could be missing something (and personally the chat works for me + my unsubscribed test account). If you email us at hi@synthetic.new though we should be able to fix anything you're running into!
- pbgcp2026 3mo agoRun it on Amazon Bedrock or GCP vertex. No problems at all.
- yieldcrv 3mo agohow much does that cost
- pbgcp2026 3mo agoThere is no markup for SOTA and Open Weight is super affordable - but most important completely private. Just try it.
- yowlingcat 3mo agoGLM5.2 is not available on bedrock or gcp vertex. They aren't real options.
- devld 3mo agoBedrock does not have GLM 5.2 and likely will not for quite some time. It seems like they are doing that on purpose due to pressure from Anthropic. DigitalOcean has it though.
- 2muchtime 3mo agoOpencode Go/Zen claim to use infrastructure based in the EU, USA and Singapore that have a 0 retention policy.
- WorldPeas 3mo agothe (imperfect) comparison having used both for planning and execution is that GLM5.2 is too jumpy and eager to do things, often to a fault (e.g. deploying/using git when it shouldn't) while sonnet 5 was much lazier than any Claude model I have used has been, not adding an addendum to a plan that I asked for, then lying that it did when asked. Looking at the analysis[0] I don't think it's worth it for me. Maybe for others. Fable was certainly much better. [0]: https://artificialanalysis.ai/models/claude-sonnet-5 https://artificialanalysis.ai/models/claude-sonnet-5
- nsoonhui 3mo agoYour benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me
- XCSme 3mo agoIt is on top for many benchmarks, only not the coding/agentic ones. Still one of the most intelligent models overall, most likely to get any question you ask correctly (without tools).
- Zababa 3mo agoNot in my experience, it tends to pick up subtle orientations given in a question (like "which is better, A and B?" and in the context you add you list a few things for A and B) and will absolutely run with them even if they're not true. Has been an issue with Gemini models since at least 3.0. Maybe that makes them great roleplaying models, but for factual information they just run with the slightest hint in one direction or another and never really push back objectively.
- BoorishBears 3mo agoThis guy had a terrible broken benchmark that gets hawked every release, and I wish HN would ban accounts that essentially exist to hawk a personally owned site, especially such a bad one.
- UqWBcuFx6NV4r 3mo agoIf you were right, the karma system would largely take care of this. It really sounds like this is more of your personal view
- BoorishBears 3mo agoKarma systems are never perfect, and most people will not assume this is a pattern. (ie. won't feel the need to downvote them just for having yet another crappy AI benchmark) I only recognize it because I build a product that leaves me looking for information on every major release... and every major release a new crop of folks reply confused about the anomalies on top of anomalies that they're seeing, and they slowly learn this person is just way more unserious than the dogged distribution would imply.