4 ms·
If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core
by gvkhna 1mo ago
If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?
- e9 1mo agoSomeone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.
- gvkhna 1mo agoWhile possible the amount of variation in serving infrastructure is unlikely to land with actually giving the exact same errors zhipu does. It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway. https://www.tomshardware.com/tech-industry/artificial-intelligence/z-ai-powers-up-1gw-ai-data-center-built-entirely-on-chinese-chips https://www.tomshardware.com/tech-industry/artificial-intell...
- vitorgrs 1mo agoThe reasoning levels are the same as GLM 5.3. GLM 5.3 is still not open... I believe it's GLM 5.3 Flash or Air.
- weiran 1mo agoReasoning levels are often just injected system prompts so not a great way to fingerprint models.
- eli 1mo agoBut it's an error, not a response.
- ggcr 1mo agoZiphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5
- LaurensBER 1mo agoThere's three options here: - The provider has a massive amount of (unused) hardware. Google or Cursor seem most likely - The model is extremely efficient, beyond anything we've seen so far - Whomever made the model has improved the cache efficiency in such a way that it's very cheap to serve. See e.g Deepseeks or Xiaomi caching (pre-price increase)
- re-thc 1mo agoOption 4: the claimed capacity is not true. Real world usage hasn’t reached anywhere close to it.
- johndough 1mo ago> 1 quadrillion tokens per day on Nous portal If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780 https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.
- SyneRyder 1mo agoI was hitting 429 overloaded regularly with Ox on OpenRouter yesterday, but a lot of that turned out to be problems with my harness. I fixed some bugs, improved the back-off, and I haven't hit a 429 error since (touch wood). OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day. https://x.com/OpenRouter/status/2091912024922177562 https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300 https://x.com/opencode/status/2090544355824038300