4 ms·
As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between pro
by XCSme 3mo ago
As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.
- yieldcrv 3mo agoWhat’s everyone favorite GLM provider? z.ai doesnt always have the most reliable AI but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh So I’ll consider all, fastest tokens/sec wins
- eli 3mo agoFireworks.ai is solid. And if you care more about speed than cost they have a "fast" variant that I think just throws more hardware at the model for about 2x the cost.
- david-gpu 3mo agoThe privacy policy indicates that they track you and share your data to ad networks like Meta. Yikes.
- pranaybhatia 3mo agoHi, PM at Fireworks here. We have zero data retention so we do not log any of your API requests. Realize you're talking about website activity which is different and will check and update on that too.
- ricardobeat 3mo agoFast variants are usually quantized to NVFP4, which incurs a slight degradation in intelligence.
- eli 3mo agoIt isn't.
- Onavo 3mo ago> the Chinese one won’t reply to subpoenas so thats a value add tbh That's not something that's definite. They are not quite like the Russians. A lot of the governments in Asia are overly pragmatic and will happily strong arm their companies to throw users under the bus for the sake of a trade deal. There's a reason why Snowden ran to the Russians and not China. Also, if they have any subsidiaries in the US, they may not have a choice in the matter.
- reissbaker 3mo agoI'm biased because I run an inference company, https://synthetic.new https://synthetic.new. That being said I think we're pretty good at serving at GLM-5.2 — and other models, like Kimi K2.7! — and our privacy policy is quite good: zero data retention for prompts and completions on API requests. Our average streaming TPS for GLM-5.2 (aka, tokens after factoring out time-to-first-token, which varies based on geography) is 97tps over the last 24hrs, although it's slightly lower at peak traffic in the mornings PST where it's 50-70 tps. We're also subscription-based which is nicer for coding than e.g. Fireworks which is per-token billing.
- yieldcrv 3mo agogot a 500 error page on the site's chat, but I'll try the API
- reissbaker 3mo agoInteresting: I don't see anything in our error logs but we could be missing something (and personally the chat works for me + my unsubscribed test account). If you email us at hi@synthetic.new though we should be able to fix anything you're running into!
- pbgcp2026 3mo agoRun it on Amazon Bedrock or GCP vertex. No problems at all.
- yieldcrv 3mo agohow much does that cost
- pbgcp2026 3mo agoThere is no markup for SOTA and Open Weight is super affordable - but most important completely private. Just try it.
- yowlingcat 3mo agoGLM5.2 is not available on bedrock or gcp vertex. They aren't real options.
- devld 3mo agoBedrock does not have GLM 5.2 and likely will not for quite some time. It seems like they are doing that on purpose due to pressure from Anthropic. DigitalOcean has it though.
- 2muchtime 3mo agoOpencode Go/Zen claim to use infrastructure based in the EU, USA and Singapore that have a 0 retention policy.