3 ms·
Very impressive speed. With a context window of 40K however, usability is limited.
by poly2it 1y ago
Very impressive speed. With a context window of 40K however, usability is limited.
- wild_egg 1y agoPost says 131k context though? What did I miss?
- mehdibl 1y agoThe PR confuse a but 32k/64k and 131k if paid API. Also this model https://huggingface.co/Qwen/Qwen3-235B-A22B https://huggingface.co/Qwen/Qwen3-235B-A22B Is native 32k. So the 64k and 131k use ROPE that is not the best for effective context. While https://qwenlm.github.io/blog/qwen3-coder/ https://qwenlm.github.io/blog/qwen3-coder/ it's 256k native https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct.
- asb 1y agoThe situation is very confusing, but the tweet that went out with the announcement indicates it's not full 131k context yet and that is coming "soon"https://xcancel.com/CerebrasSystems/status/1943765301109420285 https://xcancel.com/CerebrasSystems/status/19437653011094202...
- diggan 1y agoThe first paragraph contains: > Cerebras Systemstoday [sic] announced the launch of Qwen3-235B with full 131K context support on its inference cloud platform Then later: > Cline users can now access Cerebras Qwen models directly within the editor—starting with Qwen3-32B at 64K contexton the free tier. This rollout will expand to include Qwen3-235B with 131K context Not sure where you get the 40K number from.