3 ms·
> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window
by nh43215rgb 20d ago
> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.
> Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.
> The model will be released with open weights on October 15.
I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5).
I wonder if other Chinese labs like Kimi/Moonshot will follow suit.
- Bolwin 20d agoMoonshot has already teased K3.1 so not likely
- dannyw 20d agoK3.1 would likely be a deeper/longer post-train from K3, so that’d make sense. It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).
- JohnsonZou 20d agoAnother possible reason is that the number 4 is considered unlucky in traditional Chinese culture.
- zozbot234 19d agoParent commenter hinted at that. Yet DeepSeek has released their V4 which was hugely successful, and even their new architecture is marked V4.1. Qwen internals mark their Flash-Next model, also very compelling, as "qwen4exp". So both of them are bucking the negative stereotype.
- torginus 19d agoDunno, if US labs would embrace this silly logic, then Anthropic would be compelled to release Fable/Opus 6 instead of a .1 release
- Tepix 19d ago600b-a27b doesn’t sound enticing. Also with the higher number of active parameters compared to GLM 5.3 flash and DeepSeek V4/4.1 flash, I don’t see how they want to be more efficient at inference.
- NooneAtAll3 19d agoI kinda wish everyone just used dates instead...
- NetOpWibby 19d agoUsing dates immediately makes you look dated and everyone is in this constant race to be first!!1!!1! I agree with you though, ChronVer all the way.