6 ms·
A relief to see the Qwen team still publishing open weights, after the kneecapping [1] and departures of Junyang Lin and others [2]! [1] https://news.ycombinat
by bertili 6mo ago
A relief to see the Qwen team still publishing open weights, after the kneecapping [1] and departures of Junyang Lin and others [2]!
[1] https://news.ycombinator.com/item?id=47246746 https://news.ycombinator.com/item?id=47246746
[2] https://news.ycombinator.com/item?id=47249343 https://news.ycombinator.com/item?id=47249343
- guitcastro 6mo agoI really wish they released qwen-image 2.0 as open weights.
- zozbot234 6mo agoThis is just one model in the Qwen 3.6 series. They will most likely release the other small sizes (not much sense in keeping them proprietary) and perhaps their 122A10B size also, but the flagship 397A17B size seems to have been excluded.
- bertili 6mo agoIs there any source for these claims?
- anonova 6mo agoA Qwen research member had a poll on X asking what Qwen 3.6 sizes people wanted to see: https://x.com/ChujieZheng/status/2039909917323383036 https://x.com/ChujieZheng/status/2039909917323383036 Likely to drive engagement, but the poll excluded the large model size.
- zozbot234 6mo agohttps://x.com/ChujieZheng/status/2039909917323383036 https://x.com/ChujieZheng/status/2039909917323383036 is the pre-release poll they did. ~397B was not a listed choice and plenty of people took it as a signal that it might not be up for release.
- stingraycharles 6mo ago397A17B = 397B total weights, 17B per expert?
- wongarsu 6mo ago397B params, 17B activated at the same time Those 17B might be split among multiple experts that are activated simultaneously
- zackangelo 6mo ago17b per token. So when you’re generating a single stream of text (“decoding”) 17b parameters are active. If you’re decoding multiple streams, it will be 17b per stream (some tokens will use the same expert, so there is some overlap). When the model is ingesting the prompt (“prefilling”) it’s looking at many tokens at once, so the number of active parameters will be larger.
- littlestymaar 6mo agoThat's not how it works. Many people get confused by the “expert” naming, when in reality the key part of the original name “sparse mixture of experts” is sparse. Experts are just chunks of each layers MLP that are only partially activated by each token, there are thousands of “experts” in such a model (for Qwen3-30BA3, it was 48 layers x 128 “experts” per layer with only 8 active at each token)
- kylehotchkiss 6mo agoHow many people/hackernews can run a 397b param model at home? Probably like 20-30.
- bitbckt 6mo agoI'm running it on dual DGX Sparks.
- DoctorOetker 6mo agowhich exact model, and how many tokens per second for generation?
- manyatoms 6mo agoI'm interested in your experiences running dual
- r-w 6mo agoOpenRouter.
- parsimo2010 6mo agoIf you're running it from OpenRouter, you might as well use Qwen3.6 Plus. You don't need to be picky about a particular model size of 3.6. If you just want the 397b version to save money, just pick a cheaper model like M2.7.
- mistercheese 6mo agoYeah I think there’s benefits to third-party providers being able to run the large models and have stronger guarantees about ZDR and knowing where they are hosted! So Open Weights for even the large models we can’t personally serve on our laptops is still useful.
- stavros 6mo agoIt doesn't matter how many can run it now, it's about freedom. Having a large open weights model available allows you to do things you can't do with closed models.
- jonaustin 6mo agoAnd shout-out to Qwen if they release 122b -- Jeff Barr's original Gemma 4 tweet said they'd release a ~122b, then it got redacted :(
- canpan 6mo ago122b would be awesome. It is the largest size you can kinda run with a beefy consumer PC. I wondered about gemma stopping in the 30b category, it is already very strong. 122b might have been too close to being really useful.
- giancarlostoro 6mo ago> not much sense in keeping them proprietary Maybe for LLMs since everyone has their own competing LLM, but with Video models, Wan 2.2 did a rug pull, left a huge gap for the community that built around Wan 2.2 too, and I don't think a single open video model has come close since. Wan is at 2.7 now, and its been nearly a year since the last update.