3 ms·
128k context is not a limit of the model, that's a limit of implementation: "Context Length: 262,144 natively and extensible up to 1,000,000 tokens." https://
by gerdesj 1mo ago
128k context is not a limit of the model, that's a limit of implementation:
"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."
https://huggingface.co/Qwen/Qwen3.8-27B https://huggingface.co/Qwen/Qwen3.8-27B
- selcuka 1mo agoTPM means Tokens per Minute.
- zxexz 1mo agoGP is referring to GGP’s last paragraph. 150k t/m, yes, and 128k context.
- Aurornis 1mo agoWe're talking about the Cerebras implementation, which is limited to 128K. It's in the link.
- gerdesj 29d agoYes, and I said its a limit of the implementation and not the model.