4 ms·
Step 5 Preview: Advancing the Pareto Frontier
- Jacques2Marais 13d agoFor anyone else looking for the pricing: https://platform.stepfun.ai/docs/en/guides/pricing/details#pricing-for-multimodal-reasoning-models https://platform.stepfun.ai/docs/en/guides/pricing/details#p...
- sieve 13d agoI regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.
- ignoramous 13d ago> At $0.04/M Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview. > $8-12 per day only for cached reads Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model? Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview https://platform.stepfun.ai/docs/en/step-plan/overview
- sieve 13d agoYes, I meant the Flash version. I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting). I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.
- esafak 13d agoNot generous enough? How token efficient and fast is it compared with American models?
- conception 13d agoHuh wonder why they skipped 4?
- Alpha3031 13d agoSometimes 4 is skipped due to being considered unlucky.
- adrian_b 13d agoIn China and in places influenced by Chinese culture, due to homonymy between "4" and death.
- cynerx 13d agoSounds like death in Chinese.
- slekker 13d agoDo they rename it like in Japanese 7?
- conception 13d agoOh right right.
- bethekind 13d ago> Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story. Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.
- nh43215rgb 13d ago> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input. > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index. > The model will be released with open weights on October 15. I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5). I wonder if other Chinese labs like Kimi/Moonshot will follow suit.
- Bolwin 13d agoMoonshot has already teased K3.1 so not likely
- dannyw 13d agoK3.1 would likely be a deeper/longer post-train from K3, so that’d make sense. It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).
- JohnsonZou 13d agoAnother possible reason is that the number 4 is considered unlucky in traditional Chinese culture.
- zozbot234 13d agoParent commenter hinted at that. Yet DeepSeek has released their V4 which was hugely successful, and even their new architecture is marked V4.1. Qwen internals mark their Flash-Next model, also very compelling, as "qwen4exp". So both of them are bucking the negative stereotype.
- torginus 13d agoDunno, if US labs would embrace this silly logic, then Anthropic would be compelled to release Fable/Opus 6 instead of a .1 release
- 13d ago
- ghoshbishakh 13d agoTheir posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.
- derliebej 13d agoHow about adding a contested historical facts benchmark?
- garo-pro 13d agoIT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.
- tomComb 13d agoMost of these models, both open and closed, are now so over tuned to agentic and coding tasks that they no longer work well for general purpose. Kimi K3 is an exception to that – maybe you need that larger size todo well on a broader range of tasks.
- InsideOutSanta 13d agoGLM-5.3 and Kimi K3 are just below where I can use them to completely replace frontier models. Oddly,* SWE-2 is there for me. If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options. * Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.
- BoppreH 13d agoIn their first demo video, to make a 3D render of the photo, the thinking trace gives away the game: > Interesting! It turns out there's already an existing project here [...] The project is fully built [...] I'm always astounded how little effort is put into checking the AI answers displayed in these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos constantly showed their AIs giving wrong answers.
- BoppreH 13d agoSince people seem interested in this comment of mine, here's another fun line I noticed in the same video, at the very top of the logs, just before the Step 5 agent found the already-completed project: > Error: OpenAI API error (403): {"message":"model water18-new is not available for user i-yuliang [trace_id=bfcfdd6bcdc236ca18d009c65cca52e4 code=40004]", "type":"invalid_request_error", "param":null, "code":null} Here's a still frame for posterity (apologies for the quality, the original video is tiny): https://boppreh.com/room.jpg https://boppreh.com/room.jpg Demos are demos and recreating scenarios is to be expected, but oh boy, somebody should review these things before publishing.
- deleted 12d ago[deleted]
- Bolwin 13d agoI think it's more likely they recorded the video when the project was already done than cheated
- BoppreH 13d agoYeah, I agree that's the most likely explanation. My point is that these demo hiccups should not show up in your announcement, it casts doubt on the product for no good reason. It's sloppy.
- dofm 13d ago"Pareto frontier" really is the new "web-scale".
- wrs 13d agoBased on the examples, It’s like the thing took a writing class from Claude, but isn’t quite as smart — kind of worst of both worlds. The “Interactive Reporting” one is particularly terrible. A huge amount of waffly padding around a thesis that may or may not exist. So if the goal is to make long fancy reports that nobody will read, then, yes, very cost-effective. “the question it raises matters more than the answer” ??? “So ‘the river drifts from cool to warm’ is not a figure of speech.” Oh really?
- cboyardee 12d ago[dead]
- segmondy 12d agoThe previous Step models were pretty decent, but unfortunately for them, their model reasons too much and too long and the Qwen/Kimi/DeepSeek/GLM have been stronger. Hopefully this doesn't reason too long to get to the answer. I welcome any open model, the more the merrier.
- just60sec 11d ago[flagged]